Senior / Staff ML Ops Engineer
New
W
WaabiAutonomous Transportation
Remote US & CanadaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPythonKubernetesPyTorchCI/CDTerraformHelmDistributed Systems
Requirements
- 5+ years of software or infrastructure engineering, specifically with platforms for other engineers or ML/data-intensive systems.
- Hands-on Kubernetes expertise including GPU scheduling, autoscaling, Helm, networking, and cluster debugging.
- Excellent Python skills and a track record of designing intuitive APIs and CLIs.
- Practical AWS depth including object storage, IAM, GPU compute, networking, cost management, and Infrastructure as Code (Terraform/Pulumi).
- Experience with distributed training in PyTorch (DDP, FSDP) and experiment tracking/model registry tooling.
- Fluency with containers, CI/CD, and modern build systems for large monorepos.
- Proven ability to influence technical direction and lead migrations without formal authority.
- Strong user empathy and product instincts with a collaborative, ownership-oriented mindset.
Responsibilities
- Build and evolve training infrastructure on Kubernetes including GPU scheduling, autoscaling, and multi-node distributed jobs.
- Develop developer-facing surfaces such as CLIs, SDKs, and paved-path templates.
- Shorten inner-loop development latency and improve training run efficiency.
- Strengthen the data and artifact layer to handle high-throughput loading of large multimodal sensor data.
- Implement CI/CD for models, observability across the ML stack, and model registry/lineage systems.
- Collaborate with Security and IT to build self-serve access controls and cost governance guardrails.
- Drive internal adoption of infrastructure tools through user feedback, documentation, and support.
View Full Description & ApplyYou'll be redirected to the employer's site