Senior Site Reliability Engineer
New
C
CleraAI/ML Infrastructure
EU, UK, or North AmericaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- PythonBashGCPKubernetesGoGrafanaPrometheusTerraform
Requirements
- 3+ years of Site Reliability Engineering or production SRE experience.
- Strong proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
- Hands-on experience with Kubernetes for cluster and workload management.
- Infrastructure as Code experience using Terraform, Deployment Manager, or similar.
- Scripting and automation skills in Python, Bash, or Go.
- Experience with observability stacks including Prometheus, Grafana, OpenTelemetry, logging, and tracing.
- Proven ability to define and implement SLOs/SLIs and error budgets.
- Experience with incident management, post-incident reviews, and on-call rotations.
Responsibilities
- Design, evolve, and scale cloud infrastructure on GCP.
- Build tooling and automation to promote team autonomy and reduce toil.
- Advance the observability platform to improve MTTR and system visibility.
- Build transparency into infrastructure costs and drive cost optimization initiatives.
- Champion reliability practices including SLOs/SLIs, error budgets, and post-incident reviews.
- Lead on-call rotations and incident management.
- Partner with engineering teams to help them leverage GCP effectively and govern usage at scale.
View Full Description & ApplyYou'll be redirected to the employer's site