Senior Infrastructure Engineer — IDC / Bare-Metal Kubernetes
New
B
BinanceBlockchain, Cryptocurrency
Asia / Japan, Tokyo / Hong KongFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Fluent in English or Mandarin
- Experience
- 5+ years
- Required Skills
- PythonCiscoKubernetesGoLinuxTerraformAnsible
Requirements
- 5+ years in data center / infrastructure / platform engineering
- Hands-on experience building physical infrastructure from scratch: rack layout, network topology, server commissioning, and coordinating cross-connects and remote hands with colo / vendors
- Practical experience designing and operating OOB management networks (BMC, IPMI, Redfish)
- Have stood up production-grade self-hosted Kubernetes from scratch, and can independently debug cluster-level issues (CNI, CSI, storage)
- Strong Linux systems administration and performance tuning (kernel, networking, storage I/O)
- Bare-metal automation experience with at least one of: MAAS, Tinkerbell, Cluster API
- Proficient with Terraform, Ansible, and at least one scripting language (Python / Go / Bash)
- Experience with Cisco network and related techniques (VLAN, LACP/LAG, BGP, ACL, etc)
- Experience with Palo Alto firewall configuration
- Experience with storage systems (NetApp, Dell EMC, Pure Storage)
- Fluent in English or Mandrain
Responsibilities
- Design rack layout, network topology, and overall configuration for colocation
- Design and implement an isolated Out-of-Band (OOB) management network (BMC, IPMI, Redfish) with full security hardening
- Coordinate with colocation and hardware vendors — servers, switches, cross-connects, uplinks, remote hands, and cage/rack administration
- Build bare-metal automation for fleet-wide zero-touch OS provisioning (MAAS / Tinkerbell / Cluster API)
- Set up production-grade self-hosted Kubernetes clusters from scratch and maintain the core stack: Cilium (CNI), GitOps (ArgoCD), etc.
- Own cluster lifecycle: upgrades, scaling, node maintenance, incident response
- Write SOPs, runbooks, and post-mortems; participate in 24×7 on-call rotation
- Establish hybrid-cloud interconnect between IDC and public cloud (VPN / Direct Connect / peering)
- Pave the way for scale-out: GPU workloads (NVIDIA GPU Operator, passthrough, MIG) and VM workloads running alongside containers (KubeVirt or similar)
View Full Description & ApplyYou'll be redirected to the employer's site