ML Infra Engineer

New
H
Humble RoboticsAutonomous Vehicles
Eligible to work in the United StatesFull-Time
SalaryThis role is eligible for base salary + benefits + equity compensation.
Apply NowOpens the employer's application page

Job Details

Required Skills
CI/CDLinuxTerraformAnsibleDistributed Systems

Requirements

  • Experience building and operating high-availability web services on cloud infrastructure.
  • Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Ansible).
  • Experience building and maintaining CI/CD pipelines and managing deployments.
  • Fluent in security fundamentals including Linux hardening, network security, and cryptographic principles.
  • Hands-on experience with cluster scheduling systems for large-scale batch computation.
  • Proficiency in reading, writing, and extending non-trivial code.
  • Experience managing large, high-performance ML training clusters (preferred).
  • Knowledge of distributed training frameworks and high-performance networking for ML (preferred).
  • Prior experience at an early-stage autonomous vehicle or robotics company (preferred).

Responsibilities

  • Develop data collection infrastructure to move sensor data from vehicles to the ML platform.
  • Build batch compute pipelines for cataloging and curating training datasets.
  • Design and scale distributed ML training on GPU clusters.
  • Manage performance, observability, efficiency, and security across the full ML pipeline.
  • Partner with the ML team to identify workflow bottlenecks and build infrastructure solutions.
View Full Description & ApplyYou'll be redirected to the employer's site
This role is eligible for base salary + benefits + equity compensation.
Apply Now