Senior Site Reliability Engineer
New
C
CleraAI/ML Geospatial
Open to candidates based in EU, UK, or North America (Canada, United States, and select European countries including Denmark, Estonia, France, Netherlands, Portugal, Sweden, Switzerland, and the United Kingdom)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- PythonBashGCPKubernetesGoGrafanaPrometheusTerraform
Requirements
- 3+ years of Site Reliability Engineering or production SRE experience.
- Proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
- Hands-on experience with Kubernetes for cluster and workload management.
- Infrastructure as Code experience using tools such as Terraform or Deployment Manager.
- Scripting and automation skills in Python, Bash, or Go.
- Strong observability stack experience: Prometheus, Grafana, OpenTelemetry, logging, and distributed tracing.
- Experience with incident management, post-incident reviews, and on-call rotation.
- Ability to define and implement SLOs, SLIs, and error budgets.
Responsibilities
- Design and evolve cloud infrastructure on GCP at scale.
- Build internal tooling and automation that promote team autonomy and self-service.
- Advance the observability platform (metrics, logging, tracing) to reduce MTTR.
- Build visibility into infrastructure costs and drive governance and optimization initiatives.
- Champion reliability best practices including SLOs, SLIs, error budgets, and DORA metrics.
- Lead incident management, facilitate post-incident reviews, and participate in on-call rotation.
View Full Description & ApplyYou'll be redirected to the employer's site