Senior Software Engineer, Machine Learning Infrastructure & Automation
New
F
falGenerative AI infrastructure
Remote - USAFull-TimeSenior
Salary$170K - $230K; $170K – $230K • Offers Equity
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of software engineering background
- Required Skills
- DockerPythonPyTorchCI/CDGitHub ActionsDistributed Systems
Requirements
- Have 5+ years of software engineering background.
- Be proficient in Python and have experience building production infrastructure and developer tooling.
- Have experience designing and operating CI/CD systems using GitHub Actions or comparable technologies.
- Understand automated testing, build systems, dependency management, caching, and parallel execution.
- Have experience with containerized workloads, Docker, and cloud infrastructure.
- Be able to design reliable distributed systems and debug complex infrastructure failures.
- Understand observability, including logs, metrics, tracing, and automated alerting.
- Have a demonstrated ability to eliminate manual processes through automation.
- Have experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems.
- Be familiar with NVIDIA GPU architectures and multi-GPU environments.
- Have experience building automated inference benchmarks or ML quality evaluation frameworks.
- Experience with agentic coding tools, such as Codex or Claude Code, or building custom AI engineering agents is a plus.
Responsibilities
- Design, build, and maintain automated testing, validation, and deployment pipelines for ML models and inference services.
- Reduce CI execution times through parallelization, caching, test selection, and efficient use of compute resources.
- Develop systems to test model outputs, detect quality regressions, and validate changes across models, GPU architectures, and configurations.
- Build continuous performance testing for inference latency, throughput, GPU utilization, and cost.
- Automate validation of model pricing, billing configurations, API schemas, and deployments before production.
- Develop AI-powered automation and agentic coding systems to diagnose CI failures, identify regressions, propose fixes, and streamline engineering workflows.
- Build deployment safeguards, verification, rollback mechanisms, and monitoring for reliable model releases.
- Identify repetitive ML team tasks and build tools and systems to automate them.
View Full Description & ApplyYou'll be redirected to the employer's site