Senior Site Reliability Engineer

New
C
CleraAI/ML Geospatial
Open to candidates based in EU, UK, or North America (Canada, United States, and select European countries including Denmark, Estonia, France, Netherlands, Portugal, Sweden, Switzerland, and the United Kingdom)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
PythonBashGCPKubernetesGoGrafanaPrometheusTerraform

Requirements

  • 3+ years of Site Reliability Engineering or production SRE experience.
  • Proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
  • Hands-on experience with Kubernetes for cluster and workload management.
  • Infrastructure as Code experience using tools such as Terraform or Deployment Manager.
  • Scripting and automation skills in Python, Bash, or Go.
  • Strong observability stack experience: Prometheus, Grafana, OpenTelemetry, logging, and distributed tracing.
  • Experience with incident management, post-incident reviews, and on-call rotation.
  • Ability to define and implement SLOs, SLIs, and error budgets.

Responsibilities

  • Design and evolve cloud infrastructure on GCP at scale.
  • Build internal tooling and automation that promote team autonomy and self-service.
  • Advance the observability platform (metrics, logging, tracing) to reduce MTTR.
  • Build visibility into infrastructure costs and drive governance and optimization initiatives.
  • Champion reliability best practices including SLOs, SLIs, error budgets, and DORA metrics.
  • Lead incident management, facilitate post-incident reviews, and participate in on-call rotation.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now