Senior Site Reliability Engineer
New
J
JobgetherSecurity & IT
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 6–10 years
- Required Skills
- PythonKubernetesGoPrometheusRedisTerraformDatadogDistributed Systems
Requirements
- 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering.
- Proven experience owning systems end to end from design through production operation.
- Hands-on experience defining and operating against SLIs, SLOs, and error budgets.
- Demonstrated experience responding to high-severity production incidents.
- Strong understanding of distributed-system failure modes and cloud infrastructure fundamentals.
- Extensive experience managing infrastructure through code using Terraform.
- Strong programming skills in Go or Python.
- Experience with observability platforms such as Datadog, Prometheus, Grafana, or OpenTelemetry.
- Hands-on production experience with Redis or ElastiCache.
- Excellent written and verbal English communication skills.
- Practical experience using AI tools for engineering workflows.
Responsibilities
- Own the reliability of critical production systems end to end, including instrumentation, reliability targets, operations, and performance.
- Define and maintain SLIs, SLOs, and error budgets for key services.
- Improve alert quality and lead high-severity incident response and postmortems.
- Design secure and cost-efficient infrastructure, focusing on failure modes and resilience.
- Manage infrastructure through code using Terraform.
- Develop production software and tooling to reduce operational toil.
- Conduct chaos exercises and failure testing to establish operating limits.
- Partner with product engineering teams to ensure production readiness.
View Full Description & ApplyYou'll be redirected to the employer's site