- Design, build, and maintain automated testing, validation, and deployment pipelines for ML models and inference services.
- Reduce CI execution times through parallelization, caching, test selection, and efficient use of compute resources.
- Develop systems to test model outputs, detect quality regressions, and validate changes across models, GPU architectures, and configurations.
- Build continuous performance testing for inference latency, throughput, GPU utilization, and cost.
- Automate validation of model pricing, billing configurations, API schemas, and deployments before production.
- Develop AI-powered automation and agentic coding systems to diagnose CI failures, identify regressions, propose fixes, and streamline engineering workflows.
- Build deployment safeguards, verification, rollback mechanisms, and monitoring for reliable model releases.
- Identify repetitive ML team tasks and build tools and systems to automate them.
DockerPythonPyTorch+3 more