Site Reliability Engineer

New
D
DeepslateVoice AI
GermanyFull-TimeMiddle
Salary50,000 - 70,000 EUR per year
Apply NowOpens the employer's application page

Job Details

Languages
Fluent German
Required Skills
KubernetesCI/CDTerraformDatadog

Requirements

  • Deep hands-on experience setting up, managing, and scaling self-hosted Kubernetes clusters in production.
  • Strong experience with modern Infrastructure as Code, preferably Pulumi or deep Terraform knowledge.
  • Expertise in observability tools including Datadog and OpenTelemetry for monitoring distributed systems.
  • Proven experience with PagerDuty or similar tools and building sustainable on-call cultures.
  • Hands-on experience setting up robust integration testing and CI/CD pipelines for microservices.
  • Fluent in spoken and written German.
  • Ability to navigate the ambiguity of an early-stage startup environment.
  • Strong sense of ownership and ability to proactively identify and mitigate infrastructure risks.

Responsibilities

  • Design, build, and manage cloud infrastructure using Pulumi for reproducible and secure infrastructure.
  • Orchestrate and optimize Kubernetes clusters for compute-heavy AI workloads.
  • Implement observability and monitoring using Datadog and OpenTelemetry to identify latency and bottlenecks.
  • Manage on-call and incident response processes using PagerDuty and establish blameless post-mortems.
  • Develop and maintain highly automated integration testing and CI/CD deployment pipelines.
  • Define and monitor service-level metrics (SLAs, SLOs, SLIs) to ensure product reliability.
  • Automate manual engineering toil and harden infrastructure against security threats.
View Full Description & ApplyYou'll be redirected to the employer's site
50,000 - 70,000 EUR per year
Apply Now