Staff Software Engineer - Databases SRE
New
G
Grafana LabsCloud databases
This is a remote opportunity and we are looking for candidates from the UK, Sweden, Spain or Germany.Full-TimeStaff
SalaryIn Germany, the Base compensation range for this role is €109,709 - €131,651. Benefits include equity, bonus (if applicable) and other benefits listed here.
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years engineering experience, 4+ in SRE/CRE/production engineering.
- Required Skills
- AWSPythonGCPKubernetesAzureGoTerraformHelm
Requirements
- 8+ years of engineering experience.
- 4+ years in SRE, CRE, or production engineering; formal customer reliability engineering experience is strongly preferred.
- Strong Kubernetes experience in AWS, GCP, or Azure.
- Familiarity with infrastructure-as-code tools such as Helm, Terraform, or Jsonnet.
- Technical leadership experience, including leading projects and mentoring engineers.
- Experience operating multi-tenant systems in production.
- Strong experience designing and implementing SLOs.
- Experience with one or more programming languages, such as Go, Python, or Java.
- Experience with Linux operating system internals and knowledge of networking, cloud storage, and scaling.
- Experience participating in blame-free incident response, following up on actions, and writing post-incident reviews.
- Ability to reason about performance, scaling, and failure modes.
- Ability to partner closely with product engineering teams.
Responsibilities
- Own production reliability for high-SLA and complex customer environments.
- Define and evolve per-tenant SLOs and reliability models.
- Proactively reduce SLO burn and prevent repeat incidents.
- Design and implement automation to scale reliability practices and eliminate toil.
- Improve alert quality and reduce noisy escalations.
- Lead customer-impacting incident response, investigations, and post-incident reviews.
- Improve customer observability and design solutions for reliability and scalability.
- Develop fault-tolerant design patterns across the service lifecycle.
- Partner with product engineering squads and influence feature design, technical designs, and product roadmaps.
- Review code and design documents, and teach Site Reliability Engineering practices.
View Full Description & ApplyYou'll be redirected to the employer's site