Member of Technical Staff, Cloud Orchestration

New
I
InferactAI Infrastructure
Fully remote, worldwide., Expect regular overlap with Pacific Time for critical syncs.Full-TimeMiddle
Salary$200K - $400K; $200K – $400K • Offers Equity
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSPythonGCPKubernetesGoRustTerraformHelm

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Strong experience with Kubernetes and container orchestration at scale.
  • Experience designing and implementing custom Kubernetes operators.
  • Proficiency in Python, Rust, or Go.
  • Experience with infrastructure-as-code tools such as Terraform and Helm.
  • Experience managing GPU clusters and debugging hardware issues.
  • Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.
  • Experience with ML-specific orchestration tools such as Ray or Slurm (preferred).
  • Knowledge of GPU scheduling, multi-tenancy, and resource optimization (preferred).
  • Experience deploying inference systems on large-scale GPU clusters (1,000+ nodes) (bonus).

Responsibilities

  • Design systems for cluster management, deployment automation, and production monitoring.
  • Build the operational backbone that keeps vLLM running reliably at massive scale.
  • Ensure vLLM deployments are observable, debuggable, and recoverable.
  • Enable teams to serve AI models without friction.
  • Manage GPU clusters and debug hardware-related issues.
  • Work across AWS, GCP, Azure, and on-premise infrastructure.
View Full Description & ApplyYou'll be redirected to the employer's site
$200K - $400K; $200K – $400K • Offers Equity
Apply Now