Sr Platform Engineer, ML Infrastructure
New
B
Blue River TechnologyAI, Robotics, Agriculture
Remote in the United States.Full-TimeSenior
Salary$160,000 - $287,000/year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPythonKubeflowKubernetesAirflowSparkDistributed Systems
Requirements
- 5+ years of professional software engineering experience, with a focus on platform, infrastructure, or distributed systems.
- Strong Python engineering skills, including building production services, SDKs, automation, or platform tooling.
- Experience designing, building, and operating production platform capabilities used by multiple engineering teams.
- Understanding of ML platform architecture and the end-to-end ML lifecycle, including experimentation, distributed training, model deployment, and production operations.
- Experience building and operating applications on Kubernetes and cloud platforms (AWS preferred), with an understanding of production reliability, observability, and operational best practices.
- Strong technical judgment with the ability to independently lead complex technical initiatives from discovery through production, collaborating effectively with ML engineers, infrastructure teams, and product stakeholders.
- Experience building developer platforms, tooling, or internal services (preferred).
- Experience with workflow orchestration or distributed compute technologies such as Airflow, Kubeflow, Ray, Spark, or similar systems (preferred).
- Experience designing and optimizing distributed, GPU-intensive compute platforms (preferred).
- Experience supporting production machine learning platforms in computer vision, robotics, or similar domains (preferred).
- Demonstrated technical leadership through architecture, mentorship, or influencing technical direction across teams (preferred).
Responsibilities
- Design, build, and operate scalable ML infrastructure and platform capabilities that support the full machine learning lifecycle across cloud and on-premises environments.
- Develop developer tooling, services, and infrastructure that enable ML and engineering teams to build, deploy, and operate production systems more efficiently.
- Independently lead complex technical initiatives from problem definition and architecture through implementation, production rollout, and ongoing operational ownership.
- Make sound architectural and engineering decisions that balance near-term delivery with the platform's long-term scalability, reliability, and maintainability.
- Build reliable, scalable, easy-to-use platform capabilities that improve developer productivity, simplify operations, and help engineering teams move faster.
- Partner closely with ML engineers, infrastructure engineers, and other stakeholders to understand customer needs and translate them into effective platform solutions.
- Identify and solve challenging infrastructure and platform problems, including opportunities to improve performance, reliability, scalability, and developer experience.
- Drive adoption and continuous improvement of platform capabilities by incorporating feedback from the engineering teams that use them.
- Establish a high bar for software quality, operational excellence, and production readiness across the systems and capabilities you own.
- Deliver platform solutions with measurable engineering and business impact across multiple teams and use cases.
View Full Description & ApplyYou'll be redirected to the employer's site