Staff Site Reliability Engineer
New
J
JobgetherAI/ML Infrastructure
Based in Germany, European time zonesFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSKubernetesAzureCI/CDTerraformMLOpsDistributed Systems
Requirements
- Extensive hands-on experience in Site Reliability Engineering, Production Engineering, or a closely related infrastructure role.
- Proven experience establishing or scaling SRE practices within high-growth, complex, or highly distributed technical environments.
- Deep expertise with AWS or Azure cloud infrastructure and modern cloud-native architectures.
- Strong production experience with Kubernetes, including migration, scaling, optimization, and security hardening.
- Advanced Infrastructure-as-Code expertise using Terraform or an equivalent technology.
- Demonstrated experience designing, implementing, and optimizing end-to-end CI/CD pipelines.
- Strong knowledge of observability practices and tooling across distributed applications and infrastructure.
- Experience troubleshooting complex multi-tenant, customer-hosted, or enterprise environments.
- Experience supporting production data platforms and machine learning systems.
- Practical MLOps experience, including model deployment, monitoring, and operational lifecycle management.
- Strong understanding of distributed systems, scalability, resilience, fault tolerance, and failure modes.
- Strong communication and collaboration skills, with the ability to work effectively across engineering and business functions.
Responsibilities
- Architect, deploy, operate, and continuously improve scalable, secure production environments, with a strong preference for AWS-based infrastructure.
- Lead reliability initiatives across multiple engineering streams and establish consistent SRE practices throughout the organization.
- Design, evolve, migrate, and optimize Kubernetes-based infrastructure, including production hardening and scaling.
- Establish and enforce robust Infrastructure-as-Code standards using Terraform or equivalent technologies.
- Define, implement, and operationalize SLIs, SLOs, error budgets, and other reliability practices.
- Strengthen observability across applications, infrastructure, data pipelines, and ML systems.
- Design and optimize CI/CD pipelines across the software lifecycle.
- Lead incident response for complex, cross-system failures and drive post-incident reviews.
- Support and productionize ML workloads by implementing MLOps practices.
- Mentor engineers and collaborate with Staff Engineers/Architects to influence technical strategy.
View Full Description & ApplyYou'll be redirected to the employer's site