Senior Site Reliability Engineer
New
C
CleraGeospatial AI/ML
EU, UK, or North America (including Canada, Denmark, Estonia, France, Netherlands, Portugal, Sweden, Switzerland, and the United Kingdom)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- PythonBashGCPKubernetesGoGrafanaPrometheusTerraform
Requirements
- 3+ years of Site Reliability Engineering or production SRE experience.
- Proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
- Hands-on experience with Kubernetes for cluster and workload management.
- Infrastructure as Code experience using tools such as Terraform or Deployment Manager.
- Scripting and automation skills in Python, Bash, or Go.
- Strong observability stack experience (Prometheus, Grafana, OpenTelemetry, logging, and tracing).
- Demonstrated ability to define and implement SLOs, SLIs, and error budgets.
- Experience with incident management, post-incident reviews, and on-call rotation.
Responsibilities
- Design and evolve cloud infrastructure on GCP for scale and resilience.
- Build internal tooling and automation to promote team autonomy and developer productivity.
- Advance the observability platform including metrics, logging, tracing, and alerting to reduce MTTR.
- Build visibility into infrastructure costs and drive optimization initiatives.
- Champion reliability best practices across engineering including SLOs, SLIs, error budgets, and post-incident reviews.
- Participate in on-call rotation and lead incident management efforts.
View Full Description & ApplyYou'll be redirected to the employer's site