Lead Software Platform Engineer, MLOps
New
T
TetraScienceScientific Data / AI
United States. Cambridge, Massachusetts, United States. San Mateo, California, United StatesFull-TimeLead
Salary200,000 - 270,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of professional experience
- Required Skills
- AWSDockerPythonTypeScriptRESTful APIsCloudFormationMLOps
Requirements
- 10+ years of professional experience in software and infrastructure engineering.
- Proven track record of designing, building, and scaling distributed, cloud-native systems in production.
- Demonstrated technical leadership or architecture experience in system design and scalability.
- Deep, hands-on experience taking LLM-based systems to production (RAG, prompt/model versioning, function calling).
- Extensive experience building and maintaining production AI/ML infrastructure as a multi-tenant product.
- Expert-level coding skills in TypeScript and Python.
- Production-level experience with a model registry and serving stack, ideally Databricks MLflow.
- Proficiency in API-first design (REST, OpenAPI) and API security.
- Solid working knowledge of AWS and containerized workloads (Docker).
- Familiarity with infrastructure-as-code such as CloudFormation or CDK.
- Experience defining observability, SLI/SLO/SLA for production systems.
- Experience designing security for multi-tenant platforms, specifically handling sensitive data (PII/PHI) and LLM risks.
Responsibilities
- Architect the AI/ML platform service and API surface for customer and internal use.
- Manage the end-to-end model and prompt lifecycle using Databricks MLflow and AWS Bedrock.
- Design inference substrates for real-time and batch AI workloads, including capacity planning and concurrency control.
- Implement RAG architectures and agentic layers including tool/function calling and agent runtimes.
- Develop security guardrails, PII/PHI handling, and tenant data boundaries within the AI platform.
- Build evaluation and quality infrastructure, including regression gates and drift detection.
- Establish observability, SLI/SLO/SLA models, and audit trails for validated environments.
- Contribute to infrastructure-as-code (CloudFormation/CDK) and deployment automation.
- Lead design reviews and mentor engineers on distributed systems and AI engineering practices.
View Full Description & ApplyYou'll be redirected to the employer's site