ML Infra Engineer
New
H
Humble RoboticsAutonomous Vehicles
Eligible to work in the United StatesFull-Time
SalaryThis role is eligible for base salary + benefits + equity compensation.
Apply NowOpens the employer's application page
Job Details
- Required Skills
- CI/CDLinuxTerraformAnsibleDistributed Systems
Requirements
- Experience building and operating high-availability web services on cloud infrastructure.
- Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Ansible).
- Experience building and maintaining CI/CD pipelines and managing deployments.
- Fluent in security fundamentals including Linux hardening, network security, and cryptographic principles.
- Hands-on experience with cluster scheduling systems for large-scale batch computation.
- Proficiency in reading, writing, and extending non-trivial code.
- Experience managing large, high-performance ML training clusters (preferred).
- Knowledge of distributed training frameworks and high-performance networking for ML (preferred).
- Prior experience at an early-stage autonomous vehicle or robotics company (preferred).
Responsibilities
- Develop data collection infrastructure to move sensor data from vehicles to the ML platform.
- Build batch compute pipelines for cataloging and curating training datasets.
- Design and scale distributed ML training on GPU clusters.
- Manage performance, observability, efficiency, and security across the full ML pipeline.
- Partner with the ML team to identify workflow bottlenecks and build infrastructure solutions.
View Full Description & ApplyYou'll be redirected to the employer's site