Staff Site Reliability Engineer

New
J
JobgetherAI/ML Infrastructure
Based in Germany, European time zonesFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSKubernetesAzureCI/CDTerraformMLOpsDistributed Systems

Requirements

  • Extensive hands-on experience in Site Reliability Engineering, Production Engineering, or a closely related infrastructure role.
  • Proven experience establishing or scaling SRE practices within high-growth, complex, or highly distributed technical environments.
  • Deep expertise with AWS or Azure cloud infrastructure and modern cloud-native architectures.
  • Strong production experience with Kubernetes, including migration, scaling, optimization, and security hardening.
  • Advanced Infrastructure-as-Code expertise using Terraform or an equivalent technology.
  • Demonstrated experience designing, implementing, and optimizing end-to-end CI/CD pipelines.
  • Strong knowledge of observability practices and tooling across distributed applications and infrastructure.
  • Experience troubleshooting complex multi-tenant, customer-hosted, or enterprise environments.
  • Experience supporting production data platforms and machine learning systems.
  • Practical MLOps experience, including model deployment, monitoring, and operational lifecycle management.
  • Strong understanding of distributed systems, scalability, resilience, fault tolerance, and failure modes.
  • Strong communication and collaboration skills, with the ability to work effectively across engineering and business functions.

Responsibilities

  • Architect, deploy, operate, and continuously improve scalable, secure production environments, with a strong preference for AWS-based infrastructure.
  • Lead reliability initiatives across multiple engineering streams and establish consistent SRE practices throughout the organization.
  • Design, evolve, migrate, and optimize Kubernetes-based infrastructure, including production hardening and scaling.
  • Establish and enforce robust Infrastructure-as-Code standards using Terraform or equivalent technologies.
  • Define, implement, and operationalize SLIs, SLOs, error budgets, and other reliability practices.
  • Strengthen observability across applications, infrastructure, data pipelines, and ML systems.
  • Design and optimize CI/CD pipelines across the software lifecycle.
  • Lead incident response for complex, cross-system failures and drive post-incident reviews.
  • Support and productionize ML workloads by implementing MLOps practices.
  • Mentor engineers and collaborate with Staff Engineers/Architects to influence technical strategy.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now