Sr Platform Engineer, ML Infrastructure
New
J
JobgetherMachine Learning
Based in United StatesFull-TimeSenior
SalaryBase salary range of $160,000–$287,000 per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPythonKubeflowKubernetesMachine LearningAirflowSparkDistributed Systems
Requirements
- 5+ years of professional software engineering experience, particularly in platform engineering, infrastructure, or distributed systems.
- Strong Python engineering skills, including experience developing production services, SDKs, automation, or platform tooling.
- Proven experience designing, building, and operating production platforms used by multiple engineering teams.
- Solid understanding of ML platform architecture and the end-to-end machine learning lifecycle, including experimentation, distributed training, model deployment, and production operations.
- Experience building and operating applications on Kubernetes and cloud platforms, with AWS experience preferred.
- Strong understanding of production reliability, observability, scalability, and operational best practices.
- Strong technical judgment and the ability to independently drive complex initiatives from discovery through production.
- Experience with developer platforms, internal tooling, or services is preferred.
- Familiarity with workflow orchestration or distributed computing technologies such as Airflow, Kubeflow, Ray, or Spark.
- Experience designing or optimizing distributed, GPU-intensive compute platforms for ML training or inference.
Responsibilities
- Design, build, and operate scalable ML infrastructure and platform capabilities supporting experimentation, training, deployment, and production operations.
- Develop developer tooling, services, automation, and infrastructure that help ML and engineering teams build and operate production systems more efficiently.
- Lead complex technical initiatives independently, from problem definition and architecture through implementation, rollout, and operational ownership.
- Make architectural decisions that balance immediate delivery needs with long-term scalability, reliability, maintainability, and developer experience.
- Partner with ML engineers, infrastructure teams, and other stakeholders to understand needs and deliver effective platform solutions.
- Identify and solve challenging infrastructure problems involving performance, reliability, scalability, and operational efficiency.
- Drive adoption and continuous improvement by incorporating feedback from engineering teams using the platform.
- Maintain high standards for software quality, production readiness, observability, and operational excellence.
- Deliver platform capabilities that create measurable engineering and business impact across multiple teams and use cases.
View Full Description & ApplyYou'll be redirected to the employer's site