Senior / Staff ML Ops Engineer

New
W
WaabiAutonomous Transportation
Remote US & CanadaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPythonKubernetesPyTorchCI/CDTerraformHelmDistributed Systems

Requirements

  • 5+ years of software or infrastructure engineering, specifically with platforms for other engineers or ML/data-intensive systems.
  • Hands-on Kubernetes expertise including GPU scheduling, autoscaling, Helm, networking, and cluster debugging.
  • Excellent Python skills and a track record of designing intuitive APIs and CLIs.
  • Practical AWS depth including object storage, IAM, GPU compute, networking, cost management, and Infrastructure as Code (Terraform/Pulumi).
  • Experience with distributed training in PyTorch (DDP, FSDP) and experiment tracking/model registry tooling.
  • Fluency with containers, CI/CD, and modern build systems for large monorepos.
  • Proven ability to influence technical direction and lead migrations without formal authority.
  • Strong user empathy and product instincts with a collaborative, ownership-oriented mindset.

Responsibilities

  • Build and evolve training infrastructure on Kubernetes including GPU scheduling, autoscaling, and multi-node distributed jobs.
  • Develop developer-facing surfaces such as CLIs, SDKs, and paved-path templates.
  • Shorten inner-loop development latency and improve training run efficiency.
  • Strengthen the data and artifact layer to handle high-throughput loading of large multimodal sensor data.
  • Implement CI/CD for models, observability across the ML stack, and model registry/lineage systems.
  • Collaborate with Security and IT to build self-serve access controls and cost governance guardrails.
  • Drive internal adoption of infrastructure tools through user feedback, documentation, and support.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now