Senior Site Reliability Engineer

New
J
JobgetherSecurity & IT
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English
Experience
6–10 years
Required Skills
PythonKubernetesGoPrometheusRedisTerraformDatadogDistributed Systems

Requirements

  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering.
  • Proven experience owning systems end to end from design through production operation.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets.
  • Demonstrated experience responding to high-severity production incidents.
  • Strong understanding of distributed-system failure modes and cloud infrastructure fundamentals.
  • Extensive experience managing infrastructure through code using Terraform.
  • Strong programming skills in Go or Python.
  • Experience with observability platforms such as Datadog, Prometheus, Grafana, or OpenTelemetry.
  • Hands-on production experience with Redis or ElastiCache.
  • Excellent written and verbal English communication skills.
  • Practical experience using AI tools for engineering workflows.

Responsibilities

  • Own the reliability of critical production systems end to end, including instrumentation, reliability targets, operations, and performance.
  • Define and maintain SLIs, SLOs, and error budgets for key services.
  • Improve alert quality and lead high-severity incident response and postmortems.
  • Design secure and cost-efficient infrastructure, focusing on failure modes and resilience.
  • Manage infrastructure through code using Terraform.
  • Develop production software and tooling to reduce operational toil.
  • Conduct chaos exercises and failure testing to establish operating limits.
  • Partner with product engineering teams to ensure production readiness.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now