Senior Infrastructure Engineer — IDC / Bare-Metal Kubernetes

New
B
BinanceBlockchain, Cryptocurrency
Asia / Japan, Tokyo / Hong KongFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Fluent in English or Mandarin
Experience
5+ years
Required Skills
PythonCiscoKubernetesGoLinuxTerraformAnsible

Requirements

  • 5+ years in data center / infrastructure / platform engineering
  • Hands-on experience building physical infrastructure from scratch: rack layout, network topology, server commissioning, and coordinating cross-connects and remote hands with colo / vendors
  • Practical experience designing and operating OOB management networks (BMC, IPMI, Redfish)
  • Have stood up production-grade self-hosted Kubernetes from scratch, and can independently debug cluster-level issues (CNI, CSI, storage)
  • Strong Linux systems administration and performance tuning (kernel, networking, storage I/O)
  • Bare-metal automation experience with at least one of: MAAS, Tinkerbell, Cluster API
  • Proficient with Terraform, Ansible, and at least one scripting language (Python / Go / Bash)
  • Experience with Cisco network and related techniques (VLAN, LACP/LAG, BGP, ACL, etc)
  • Experience with Palo Alto firewall configuration
  • Experience with storage systems (NetApp, Dell EMC, Pure Storage)
  • Fluent in English or Mandrain

Responsibilities

  • Design rack layout, network topology, and overall configuration for colocation
  • Design and implement an isolated Out-of-Band (OOB) management network (BMC, IPMI, Redfish) with full security hardening
  • Coordinate with colocation and hardware vendors — servers, switches, cross-connects, uplinks, remote hands, and cage/rack administration
  • Build bare-metal automation for fleet-wide zero-touch OS provisioning (MAAS / Tinkerbell / Cluster API)
  • Set up production-grade self-hosted Kubernetes clusters from scratch and maintain the core stack: Cilium (CNI), GitOps (ArgoCD), etc.
  • Own cluster lifecycle: upgrades, scaling, node maintenance, incident response
  • Write SOPs, runbooks, and post-mortems; participate in 24×7 on-call rotation
  • Establish hybrid-cloud interconnect between IDC and public cloud (VPN / Direct Connect / peering)
  • Pave the way for scale-out: GPU workloads (NVIDIA GPU Operator, passthrough, MIG) and VM workloads running alongside containers (KubeVirt or similar)
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now