Senior Software Engineer - Databases, SRE
New
G
Grafana LabsSoftware Observability
Applicants based in Canadian timezones at this time, Canadian timezonesFull-TimeSenior
SalaryCAD 164,490 - CAD 197,389
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years engineering experience, 3+ in SRE/CRE/production engineering
- Required Skills
- PythonJavaKubernetesGoLinuxTerraformHelm
Requirements
- 6+ years of engineering experience, with 3+ years in SRE, CRE, or production engineering.
- Strong preference for formal customer reliability engineering experience.
- Strong Kubernetes experience within AWS, GCP, or Azure environments.
- Familiarity with infrastructure-as-code tooling such as Helm, Terraform, and Jsonnet.
- Experience operating multi-tenant systems in production environments.
- Proven experience designing and implementing Service Level Objectives (SLOs).
- Proficiency in one or more programming languages (e.g., Go, Python, Java).
- Solid understanding of Linux operating system internals.
- Knowledge of networking, cloud storage, and scaling concepts.
- Proven ability to participate in blame-free incident response and write high-quality PIRs.
- Ability to reason effectively about performance, scaling, and failure modes.
Responsibilities
- Partner closely with product engineering squads in an embedded model.
- Own production reliability for high-SLA and complex customer environments.
- Design and implement automation to scale reliability practices and eliminate toil.
- Define and evolve per-tenant SLOs and reliability models while ensuring targets are met.
- Lead customer-impacting incident response, post-incident reviews, and serve as a primary escalation point.
- Contribute to design docs and code reviews to ensure production scalability and operability.
- Improve alert quality and proactively reduce SLO burn to prevent repeat incidents.
View Full Description & ApplyYou'll be redirected to the employer's site