Sr Platform Engineer, ML Infrastructure

New
B
Blue River TechnologyAI, Robotics, Agriculture
Remote in the United States.Full-TimeSenior
Salary$160,000 - $287,000/year
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPythonKubeflowKubernetesAirflowSparkDistributed Systems

Requirements

  • 5+ years of professional software engineering experience, with a focus on platform, infrastructure, or distributed systems.
  • Strong Python engineering skills, including building production services, SDKs, automation, or platform tooling.
  • Experience designing, building, and operating production platform capabilities used by multiple engineering teams.
  • Understanding of ML platform architecture and the end-to-end ML lifecycle, including experimentation, distributed training, model deployment, and production operations.
  • Experience building and operating applications on Kubernetes and cloud platforms (AWS preferred), with an understanding of production reliability, observability, and operational best practices.
  • Strong technical judgment with the ability to independently lead complex technical initiatives from discovery through production, collaborating effectively with ML engineers, infrastructure teams, and product stakeholders.
  • Experience building developer platforms, tooling, or internal services (preferred).
  • Experience with workflow orchestration or distributed compute technologies such as Airflow, Kubeflow, Ray, Spark, or similar systems (preferred).
  • Experience designing and optimizing distributed, GPU-intensive compute platforms (preferred).
  • Experience supporting production machine learning platforms in computer vision, robotics, or similar domains (preferred).
  • Demonstrated technical leadership through architecture, mentorship, or influencing technical direction across teams (preferred).

Responsibilities

  • Design, build, and operate scalable ML infrastructure and platform capabilities that support the full machine learning lifecycle across cloud and on-premises environments.
  • Develop developer tooling, services, and infrastructure that enable ML and engineering teams to build, deploy, and operate production systems more efficiently.
  • Independently lead complex technical initiatives from problem definition and architecture through implementation, production rollout, and ongoing operational ownership.
  • Make sound architectural and engineering decisions that balance near-term delivery with the platform's long-term scalability, reliability, and maintainability.
  • Build reliable, scalable, easy-to-use platform capabilities that improve developer productivity, simplify operations, and help engineering teams move faster.
  • Partner closely with ML engineers, infrastructure engineers, and other stakeholders to understand customer needs and translate them into effective platform solutions.
  • Identify and solve challenging infrastructure and platform problems, including opportunities to improve performance, reliability, scalability, and developer experience.
  • Drive adoption and continuous improvement of platform capabilities by incorporating feedback from the engineering teams that use them.
  • Establish a high bar for software quality, operational excellence, and production readiness across the systems and capabilities you own.
  • Deliver platform solutions with measurable engineering and business impact across multiple teams and use cases.
View Full Description & ApplyYou'll be redirected to the employer's site
$160,000 - $287,000/year
Apply Now