Senior Backend Engineer: Machine Learning Infrastructure
New
J
JobgetherMachine learning infrastructure
Based in ItalyFull-TimeSenior
SalaryCompetitive base salary of $80,000–$120,000 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of professional experience in backend engineering, platform engineering, or a closely related discipline.
- Required Skills
- AWSPythonGCPKubernetesAzureRESTful APIsTerraformDistributed Systems
Requirements
- Have 5+ years of professional experience in backend engineering, platform engineering, or a closely related discipline.
- Bring extensive professional experience with Python.
- Have hands-on experience developing and operating software on AWS, GCP, Azure, or self-managed Kubernetes; AWS experience is particularly relevant.
- Have strong experience designing and building distributed, high-load services and APIs.
- Understand data structures, algorithms, and the trade-offs involved in selecting and implementing them.
- Be able to design solutions before implementation and understand the architectural implications of technical decisions.
- Have experience with modern AI-assisted coding and development tools and understand their appropriate use and limitations.
- Demonstrate ownership, initiative, accountability, and a proactive approach to solving problems.
- Have strong communication and collaboration skills and a supportive approach to working with teammates and cross-functional partners.
- Experience with Rust, C, C++, or Go is an advantage.
- Experience developing or contributing to ML platforms or machine learning infrastructure is valued.
- Experience with vector databases such as Qdrant, Milvus, Weaviate, OpenSearch, or pgvector is a plus.
- Experience with model serving or inference infrastructure, including LLM workloads, is advantageous.
- Experience with Infrastructure as Code tools such as Terraform is a plus.
Responsibilities
- Design, build, deploy, and operate high-load distributed backend services and APIs for machine learning infrastructure.
- Own core ML services and data pipelines from system design and implementation through deployment, observability, maintenance, and improvement.
- Build reliable, scalable, reusable infrastructure components for machine learning workloads.
- Partner with ML and product engineers to translate requirements into platform capabilities and services.
- Evaluate architectural options and trade-offs, including scalability, reliability, and operational requirements.
- Maintain service reliability, performance, observability, and maintainability.
- Identify and resolve technical and operational problems.
- Use AI-assisted development tools while retaining ownership of system design, engineering decisions, and code quality.
- Share knowledge and support teammates and cross-functional partners.
View Full Description & ApplyYou'll be redirected to the employer's site