Senior Infrastructure Engineer - GPU Compute

New
B
BoundlessAI Compute
We are a global team, and applicants from around the world are welcome to apply.Full-TimeSenior
SalaryCompetitive salary + equity allocation
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
DockerPythonBashKubernetesTypeScriptGoLinuxTerraformAnsible

Requirements

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems.
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale.
  • Strong Linux systems administration skills.
  • Proficiency in infrastructure-as-code tools such as Terraform, Ansible, or Pulumi.
  • Track record of managing mission-critical, high-throughput systems.
  • Proficiency in at least one common scripting or programming language (e.g., Python, Bash, TypeScript, Go).
  • Must include a public GitHub profile with at least 1 year of activity.
  • Comfort navigating ambiguity with a strong bias for action.
  • Strong infrastructure-as-code background in heterogeneous environments.

Responsibilities

  • Operate a heterogeneous, multi-region GPU fleet using tools like SkyPilot, Kubernetes/k3s, and cloud or on-prem providers.
  • Maximize GPU utilization and workload placement across spot, on-prem, and cloud capacity.
  • Perform bare-metal and GPU optimization, including kernel tuning, memory configuration, and network topology analysis.
  • Build secure fleet access, observability, alerting, and zero-downtime deployment patterns.
  • Drive down cost per GPU-hour through intelligent scheduling and instance management.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive salary + equity allocation
Apply Now