Senior Site Reliability Engineer

New
C
CleraAI/ML Infrastructure
Open to candidates based in the EU, UK, or North AmericaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
PythonBashGCPKubernetesGoGrafanaPrometheusTerraform

Requirements

  • 3+ years of Site Reliability Engineering or production SRE experience.
  • Strong proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
  • Hands-on experience with Kubernetes for cluster and workload management.
  • Infrastructure as Code experience such as Terraform or Deployment Manager.
  • Scripting and automation skills in Python, Bash, or Go.
  • Strong observability stack experience: Prometheus, Grafana, OpenTelemetry, logging, and tracing.
  • Proven ability to define and implement SLOs/SLIs and error budgets.
  • Experience with incident management, post-incident reviews, and on-call rotations.

Responsibilities

  • Design, evolve, and scale cloud infrastructure on GCP.
  • Build tooling and automation to promote team autonomy and reduce toil.
  • Advance observability platforms to improve MTTR and system visibility.
  • Build transparency into infrastructure costs and drive cost optimization.
  • Champion reliability practices such as SLOs/SLIs, error budgets, and post-incident reviews.
  • Lead on-call rotations and incident management in a blameless culture.
  • Collaborate with engineering teams to govern GCP usage at scale.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now