Lead Software Platform Engineer, MLOps

New
T
TetraScienceScientific Data / AI
United States. Cambridge, Massachusetts, United States. San Mateo, California, United StatesFull-TimeLead
Salary200,000 - 270,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
10+ years of professional experience
Required Skills
AWSDockerPythonTypeScriptRESTful APIsCloudFormationMLOps

Requirements

  • 10+ years of professional experience in software and infrastructure engineering.
  • Proven track record of designing, building, and scaling distributed, cloud-native systems in production.
  • Demonstrated technical leadership or architecture experience in system design and scalability.
  • Deep, hands-on experience taking LLM-based systems to production (RAG, prompt/model versioning, function calling).
  • Extensive experience building and maintaining production AI/ML infrastructure as a multi-tenant product.
  • Expert-level coding skills in TypeScript and Python.
  • Production-level experience with a model registry and serving stack, ideally Databricks MLflow.
  • Proficiency in API-first design (REST, OpenAPI) and API security.
  • Solid working knowledge of AWS and containerized workloads (Docker).
  • Familiarity with infrastructure-as-code such as CloudFormation or CDK.
  • Experience defining observability, SLI/SLO/SLA for production systems.
  • Experience designing security for multi-tenant platforms, specifically handling sensitive data (PII/PHI) and LLM risks.

Responsibilities

  • Architect the AI/ML platform service and API surface for customer and internal use.
  • Manage the end-to-end model and prompt lifecycle using Databricks MLflow and AWS Bedrock.
  • Design inference substrates for real-time and batch AI workloads, including capacity planning and concurrency control.
  • Implement RAG architectures and agentic layers including tool/function calling and agent runtimes.
  • Develop security guardrails, PII/PHI handling, and tenant data boundaries within the AI platform.
  • Build evaluation and quality infrastructure, including regression gates and drift detection.
  • Establish observability, SLI/SLO/SLA models, and audit trails for validated environments.
  • Contribute to infrastructure-as-code (CloudFormation/CDK) and deployment automation.
  • Lead design reviews and mentor engineers on distributed systems and AI engineering practices.
View Full Description & ApplyYou'll be redirected to the employer's site
200,000 - 270,000 USD per year
Apply Now