Senior Backend Engineer: Machine Learning Infrastructure
New
J
JobgetherMachine learning infrastructure
GermanyFull-TimeSenior
Salary$80,000–$120,000 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of professional experience in backend engineering, platform engineering, or a closely related discipline.
- Required Skills
- PythonKubernetesTerraformDistributed Systems
Requirements
- Have 5+ years of professional experience in backend engineering, platform engineering, or a closely related discipline.
- Bring extensive professional experience with Python.
- Have hands-on experience developing and operating software on AWS, GCP, or Azure, or working with self-managed Kubernetes; AWS experience is particularly relevant.
- Have strong experience designing and building distributed, high-load services and APIs.
- Understand data structures, algorithms, and their implementation trade-offs.
- Be able to design solutions before implementation and understand the architectural implications of technical decisions.
- Have experience with modern AI-assisted coding and development tools and understand their appropriate use and limitations.
- Bring ownership, initiative, accountability, and a proactive approach to problem-solving.
- Have strong communication and collaboration skills.
- Experience with Rust, C, C++, or Go is an advantage.
- Machine learning platform or infrastructure experience is highly valued.
- Experience operating vector databases such as Qdrant, Milvus, Weaviate, OpenSearch, or pgvector is a plus.
- Model serving or inference infrastructure experience, including LLM workloads, is advantageous.
- Experience with Infrastructure as Code tools such as Terraform is a plus.
Responsibilities
- Design, build, deploy, and operate high-load distributed backend services and APIs for machine learning infrastructure.
- Own core ML services and data pipelines from system design and implementation through deployment, observability, maintenance, and improvement.
- Build reliable, scalable, reusable infrastructure components for machine learning workloads.
- Partner with ML and product engineers to translate requirements into platform capabilities and services.
- Evaluate architectural options and trade-offs, and communicate technical decisions.
- Maintain service reliability, performance, observability, and maintainability.
- Identify and resolve technical and operational problems.
- Use AI-assisted development tools while maintaining ownership of system design, engineering decisions, and code quality.
- Share knowledge and support teammates and cross-functional partners.
View Full Description & ApplyYou'll be redirected to the employer's site