Serve as the technical owner of Tempest, Domino's scale and reliability platform.
Diagnose and drive resolution of performance bottlenecks and resource misconfigurations surfaced by scale testing.
Deliver accurate, data-driven sizing recommendations for customer-facing documentation based on empirical testing.
Strengthen observability by improving Prometheus and New Relic instrumentation.
Establish and operationalize scale testing on cloud platforms.
Partner with platform teams to enable scale and reliability testing across additional cloud providers.
Build infrastructure automation to increase team efficiency.
PythonKubernetesGrafana+1 more
Showing 1 of 13 positions
About Domino Data Lab
Domino Data Lab empowers the world's largest, AI-driven enterprises to build, deploy, and manage AI at scale. Their Enterprise AI Platform acts as a central system of record for data science teams, integrating tools, compute, and data. You will help accelerate breakthroughs for customers like Johnson & Johnson, Dell Technologies, and the US Navy. They trust Domino to industrialize AI, driving innovation while managing costs and ensuring compliance. Founded in 2013, Domino Data Lab has raised $224 million across 9 funding rounds, backed by leading investors including Sequoia Capital and NVIDIA.
How We Work
Domino Data Lab embraces a hybrid work model, blending in-office collaboration with remote flexibility. They actively support work-from-home requests, offering a $500 stipend for new employees to set up their home offices. Job postings show openings for "Remote US" and "Remote Argentina." They foster a culture that values a growth mindset and continuous improvement, believing that everything is a work in progress. Expect an environment of teaching and learning, where you are equipped with tools for success. They champion diversity, encouraging applicants from all backgrounds, genders, ethnicities, abilities, and sexual orientations.
Engineering at Domino Data Lab
Engineers at Domino Data Lab solve complex challenges in AI and MLOps. They build internal AI-assisted reliability tooling that analyzes tickets, logs, traces, and documentation to resolve outages faster. You will modernize CI/CD pipelines, expanding automation coverage and integrating AI-assisted coding tools like Claude, Codex, or Cursor. Projects span scalable back-end systems in distributed computing environments, leveraging Kubernetes, cloud platforms (AWS, Azure, GCP), and observability stacks like Prometheus and New Relic. You will also work with ML model deployment, registries, versioning, and lifecycle management tools, impacting how enterprises develop and operationalize AI.
Why Join Us
Pioneer AI industrialization for Fortune 100 companies, including Johnson & Johnson and the US Navy.
Shape the future of AI/ML platforms by building cutting-edge, AI-assisted tooling and infrastructure.
Collaborate in a hybrid, growth-minded environment that prioritizes continuous learning and diversity.
Work on high-impact projects that transform data science teams' productivity and compliance.
Receive comprehensive benefits, including equity, robust healthcare, and professional development stipends.
Benefits & Perks
Company equity (stock options vesting over four years with a one-year cliff)
401(k) or pension plan
Premium medical, dental, and vision insurance (free option for employees and families)
Flexible Paid Time Off (PTO) and paid holidays/sick time