Sr Platform Engineer, ML Infrastructure

New
J
JobgetherMachine Learning
Based in United StatesFull-TimeSenior
SalaryBase salary range of $160,000–$287,000 per year
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPythonKubeflowKubernetesMachine LearningAirflowSparkDistributed Systems

Requirements

  • 5+ years of professional software engineering experience, particularly in platform engineering, infrastructure, or distributed systems.
  • Strong Python engineering skills, including experience developing production services, SDKs, automation, or platform tooling.
  • Proven experience designing, building, and operating production platforms used by multiple engineering teams.
  • Solid understanding of ML platform architecture and the end-to-end machine learning lifecycle, including experimentation, distributed training, model deployment, and production operations.
  • Experience building and operating applications on Kubernetes and cloud platforms, with AWS experience preferred.
  • Strong understanding of production reliability, observability, scalability, and operational best practices.
  • Strong technical judgment and the ability to independently drive complex initiatives from discovery through production.
  • Experience with developer platforms, internal tooling, or services is preferred.
  • Familiarity with workflow orchestration or distributed computing technologies such as Airflow, Kubeflow, Ray, or Spark.
  • Experience designing or optimizing distributed, GPU-intensive compute platforms for ML training or inference.

Responsibilities

  • Design, build, and operate scalable ML infrastructure and platform capabilities supporting experimentation, training, deployment, and production operations.
  • Develop developer tooling, services, automation, and infrastructure that help ML and engineering teams build and operate production systems more efficiently.
  • Lead complex technical initiatives independently, from problem definition and architecture through implementation, rollout, and operational ownership.
  • Make architectural decisions that balance immediate delivery needs with long-term scalability, reliability, maintainability, and developer experience.
  • Partner with ML engineers, infrastructure teams, and other stakeholders to understand needs and deliver effective platform solutions.
  • Identify and solve challenging infrastructure problems involving performance, reliability, scalability, and operational efficiency.
  • Drive adoption and continuous improvement by incorporating feedback from engineering teams using the platform.
  • Maintain high standards for software quality, production readiness, observability, and operational excellence.
  • Deliver platform capabilities that create measurable engineering and business impact across multiple teams and use cases.
View Full Description & ApplyYou'll be redirected to the employer's site
Base salary range of $160,000–$287,000 per year
Apply Now