Senior Site Reliability Engineer

New
C
CleraGeospatial AI/ML
EU, UK, or North America (including Canada, Denmark, Estonia, France, Netherlands, Portugal, Sweden, Switzerland, and the United Kingdom)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
PythonBashGCPKubernetesGoGrafanaPrometheusTerraform

Requirements

  • 3+ years of Site Reliability Engineering or production SRE experience.
  • Proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
  • Hands-on experience with Kubernetes for cluster and workload management.
  • Infrastructure as Code experience using tools such as Terraform or Deployment Manager.
  • Scripting and automation skills in Python, Bash, or Go.
  • Strong observability stack experience (Prometheus, Grafana, OpenTelemetry, logging, and tracing).
  • Demonstrated ability to define and implement SLOs, SLIs, and error budgets.
  • Experience with incident management, post-incident reviews, and on-call rotation.

Responsibilities

  • Design and evolve cloud infrastructure on GCP for scale and resilience.
  • Build internal tooling and automation to promote team autonomy and developer productivity.
  • Advance the observability platform including metrics, logging, tracing, and alerting to reduce MTTR.
  • Build visibility into infrastructure costs and drive optimization initiatives.
  • Champion reliability best practices across engineering including SLOs, SLIs, error budgets, and post-incident reviews.
  • Participate in on-call rotation and lead incident management efforts.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now