Senior Infrastructure Engineer - GPU Compute
New
B
BoundlessAI Compute
We are a global team, and applicants from around the world are welcome to apply.Full-TimeSenior
SalaryCompetitive salary + equity allocation
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- DockerPythonBashKubernetesTypeScriptGoLinuxTerraformAnsible
Requirements
- 5+ years of infrastructure/DevOps experience operating large-scale production systems.
- Deep expertise in Kubernetes, Docker, and container orchestration at scale.
- Strong Linux systems administration skills.
- Proficiency in infrastructure-as-code tools such as Terraform, Ansible, or Pulumi.
- Track record of managing mission-critical, high-throughput systems.
- Proficiency in at least one common scripting or programming language (e.g., Python, Bash, TypeScript, Go).
- Must include a public GitHub profile with at least 1 year of activity.
- Comfort navigating ambiguity with a strong bias for action.
- Strong infrastructure-as-code background in heterogeneous environments.
Responsibilities
- Operate a heterogeneous, multi-region GPU fleet using tools like SkyPilot, Kubernetes/k3s, and cloud or on-prem providers.
- Maximize GPU utilization and workload placement across spot, on-prem, and cloud capacity.
- Perform bare-metal and GPU optimization, including kernel tuning, memory configuration, and network topology analysis.
- Build secure fleet access, observability, alerting, and zero-downtime deployment patterns.
- Drive down cost per GPU-hour through intelligent scheduling and instance management.
View Full Description & ApplyYou'll be redirected to the employer's site