Senior Site Reliability Engineer

New
C
CleraAI/ML Infrastructure
EU, UK, or North AmericaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
PythonBashGCPKubernetesGoGrafanaPrometheusTerraform

Requirements

  • 3+ years of Site Reliability Engineering or production SRE experience.
  • Strong proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
  • Hands-on experience with Kubernetes for cluster and workload management.
  • Infrastructure as Code experience using Terraform, Deployment Manager, or similar.
  • Scripting and automation skills in Python, Bash, or Go.
  • Experience with observability stacks including Prometheus, Grafana, OpenTelemetry, logging, and tracing.
  • Proven ability to define and implement SLOs/SLIs and error budgets.
  • Experience with incident management, post-incident reviews, and on-call rotations.

Responsibilities

  • Design, evolve, and scale cloud infrastructure on GCP.
  • Build tooling and automation to promote team autonomy and reduce toil.
  • Advance the observability platform to improve MTTR and system visibility.
  • Build transparency into infrastructure costs and drive cost optimization initiatives.
  • Champion reliability practices including SLOs/SLIs, error budgets, and post-incident reviews.
  • Lead on-call rotations and incident management.
  • Partner with engineering teams to help them leverage GCP effectively and govern usage at scale.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now