Site Reliability Engineer
New
D
DeepslateVoice AI
GermanyFull-TimeMiddle
Salary50,000 - 70,000 EUR per year
Apply NowOpens the employer's application page
Job Details
- Languages
- Fluent German
- Required Skills
- KubernetesCI/CDTerraformDatadog
Requirements
- Deep hands-on experience setting up, managing, and scaling self-hosted Kubernetes clusters in production.
- Strong experience with modern Infrastructure as Code, preferably Pulumi or deep Terraform knowledge.
- Expertise in observability tools including Datadog and OpenTelemetry for monitoring distributed systems.
- Proven experience with PagerDuty or similar tools and building sustainable on-call cultures.
- Hands-on experience setting up robust integration testing and CI/CD pipelines for microservices.
- Fluent in spoken and written German.
- Ability to navigate the ambiguity of an early-stage startup environment.
- Strong sense of ownership and ability to proactively identify and mitigate infrastructure risks.
Responsibilities
- Design, build, and manage cloud infrastructure using Pulumi for reproducible and secure infrastructure.
- Orchestrate and optimize Kubernetes clusters for compute-heavy AI workloads.
- Implement observability and monitoring using Datadog and OpenTelemetry to identify latency and bottlenecks.
- Manage on-call and incident response processes using PagerDuty and establish blameless post-mortems.
- Develop and maintain highly automated integration testing and CI/CD deployment pipelines.
- Define and monitor service-level metrics (SLAs, SLOs, SLIs) to ensure product reliability.
- Automate manual engineering toil and harden infrastructure against security threats.
View Full Description & ApplyYou'll be redirected to the employer's site