Site Reliability Engineering Team Lead (Principal SRE)

J
JobgetherCloud-native AI
United StatesFull-TimePrincipal
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English
Experience
8+ years
Required Skills
PythonKubernetesAzureGoCI/CDLinuxTerraform

Requirements

  • 8+ years of experience in SRE, DevOps, or cloud platforms, including team leadership or technical function ownership.
  • Demonstrated ability to establish technical direction and influence teams without direct reports.
  • Hands-on experience with Kubernetes, Docker, and Istio.
  • Strong experience with Azure (exposure to AWS and Google Cloud).
  • Experience with observability tools such as Zabbix, Prometheus, and Grafana.
  • Experience designing CI/CD pipelines and IaC solutions using Terraform and Flux.
  • Proficiency in Python, Go, or Shell scripting.
  • Strong UNIX/Linux expertise including configuration, troubleshooting, and networking (DNS, HTTP/S, TLS).
  • Deep understanding of high-availability architecture, failover, and redundancy.
  • Excellent written and verbal communication skills in English.

Responsibilities

  • Provide technical leadership across the SRE function, mentoring and developing engineers across multiple locations.
  • Establish technical direction, priorities, and engineering standards for reliability practices.
  • Own and execute a reliability roadmap covering a 2–3 quarter planning horizon.
  • Define and govern SLI, SLO, and SLA frameworks to meet availability targets of up to 99.95%.
  • Design and maintain a sustainable on-call model and lead incident response as a Tier 2 escalation point.
  • Lead Production Readiness reviews and drive root cause analysis for production incidents.
  • Drive CI/CD automation for deployments, rollbacks, and operational processes.
  • Partner with DevOps and architecture teams to integrate reliability principles into the SDLC.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now