Senior Site Reliability Engineer
New
C
CleraAI/ML Infrastructure
Open to candidates based in the EU, UK, or North AmericaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- PythonBashGCPKubernetesGoGrafanaPrometheusTerraform
Requirements
- 3+ years of Site Reliability Engineering or production SRE experience.
- Strong proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
- Hands-on experience with Kubernetes for cluster and workload management.
- Infrastructure as Code experience such as Terraform or Deployment Manager.
- Scripting and automation skills in Python, Bash, or Go.
- Strong observability stack experience: Prometheus, Grafana, OpenTelemetry, logging, and tracing.
- Proven ability to define and implement SLOs/SLIs and error budgets.
- Experience with incident management, post-incident reviews, and on-call rotations.
Responsibilities
- Design, evolve, and scale cloud infrastructure on GCP.
- Build tooling and automation to promote team autonomy and reduce toil.
- Advance observability platforms to improve MTTR and system visibility.
- Build transparency into infrastructure costs and drive cost optimization.
- Champion reliability practices such as SLOs/SLIs, error budgets, and post-incident reviews.
- Lead on-call rotations and incident management in a blameless culture.
- Collaborate with engineering teams to govern GCP usage at scale.
View Full Description & ApplyYou'll be redirected to the employer's site