Site Reliability Engineering Team Lead (Principal SRE)
J
JobgetherCloud-native AI
United StatesFull-TimePrincipal
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 8+ years
- Required Skills
- PythonKubernetesAzureGoCI/CDLinuxTerraform
Requirements
- 8+ years of experience in SRE, DevOps, or cloud platforms, including team leadership or technical function ownership.
- Demonstrated ability to establish technical direction and influence teams without direct reports.
- Hands-on experience with Kubernetes, Docker, and Istio.
- Strong experience with Azure (exposure to AWS and Google Cloud).
- Experience with observability tools such as Zabbix, Prometheus, and Grafana.
- Experience designing CI/CD pipelines and IaC solutions using Terraform and Flux.
- Proficiency in Python, Go, or Shell scripting.
- Strong UNIX/Linux expertise including configuration, troubleshooting, and networking (DNS, HTTP/S, TLS).
- Deep understanding of high-availability architecture, failover, and redundancy.
- Excellent written and verbal communication skills in English.
Responsibilities
- Provide technical leadership across the SRE function, mentoring and developing engineers across multiple locations.
- Establish technical direction, priorities, and engineering standards for reliability practices.
- Own and execute a reliability roadmap covering a 2–3 quarter planning horizon.
- Define and govern SLI, SLO, and SLA frameworks to meet availability targets of up to 99.95%.
- Design and maintain a sustainable on-call model and lead incident response as a Tier 2 escalation point.
- Lead Production Readiness reviews and drive root cause analysis for production incidents.
- Drive CI/CD automation for deployments, rollbacks, and operational processes.
- Partner with DevOps and architecture teams to integrate reliability principles into the SDLC.
View Full Description & ApplyYou'll be redirected to the employer's site