Staff Site Reliability Engineer
New
J
JobgetherSecurity & IT
Based in IndiaFull-TimeSenior
Salary177,000 - 240,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role
- Required Skills
- AWSKubernetesTypeScriptGoDatadogDistributed Systems
Requirements
- 10+ years of engineering experience.
- 3+ years in SRE, production engineering, or a reliability-focused Staff Engineer role.
- Experience owning reliability at a platform or organizational level.
- Deep practical experience with SLIs, SLOs, and error budgets.
- Strong incident leadership experience and management of high-severity incidents.
- Advanced understanding of distributed-system failure modes.
- Hands-on experience with Kubernetes, AWS, and modern observability platforms (e.g., Datadog).
- Proficiency in Go, TypeScript, or a similar language.
- Ability to influence teams without direct authority and drive organizational change.
- Strong coaching and mentoring skills.
- Practical experience using AI tools for incident investigation and engineering tooling.
- Exceptional written and verbal communication skills.
Responsibilities
- Define and implement SLIs and SLOs for critical production request paths.
- Introduce and champion error budgets as a framework for balancing reliability and delivery.
- Strengthen the incident management lifecycle from detection through postmortems.
- Collaborate with infrastructure teams to improve alert quality and operational tooling.
- Lead reliability assessments for high-risk changes and new services.
- Conduct deliberate failure testing, game days, and chaos exercises.
- Coach Staff and Lead engineers to foster distributed SRE ownership.
- Develop operational standards for runbooks, on-call practices, and change safety.
- Promote the effective use of AI for incident investigation and observability.
- Contribute production fixes directly through code and infrastructure changes.
View Full Description & ApplyYou'll be redirected to the employer's site