Staff Site Reliability Engineer
A
ArcadiaEnergy Intelligence
Chennai, Tamil Nadu, India, Must work in the India timezone and collaborate daily with US-based SRE leadership.Full-TimeStaff
SalaryCompetitive compensation based on market standards.
Apply NowOpens the employer's application page
Job Details
- Experience
- 8–14 years of experience in SRE/DevOps/Cloud Engineering
- Required Skills
- AWSJenkinsKubernetesGrafanaPrometheusCI/CDTerraformGitHub ActionsCloudFormation
Requirements
- 8–14 years of experience in SRE/DevOps/Cloud Engineering.
- Deep, hands-on expertise with AWS (EKS, IAM, RDS, EC2, VPC, CloudWatch, Lambda, SQS).
- Strong Terraform skills with experience managing complex, multi-environment state.
- Advanced Kubernetes knowledge, including CNI troubleshooting and cluster upgrades.
- Experience with CI/CD pipeline design (Jenkins/Groovy, GitHub Actions, ArgoCD, or FluxCD).
- Observability stack experience with Prometheus, Grafana, Datadog, or equivalent.
- Proven mentorship ability and technical leadership experience.
- Strong written and verbal communication skills for daily collaboration with US teams.
- Automation-first mindset with a track record of reducing operational toil.
- Solid incident management experience in production environments.
- Ability to operate with autonomy under ambiguity.
Responsibilities
- Own and deliver SRE projects end-to-end from scoping through implementation and documentation.
- Serve as a technical anchor by conducting design reviews, debugging complex issues, and mentoring engineers.
- Design and implement infrastructure solutions across AWS (EKS, VPC, RDS, IAM) using Terraform and CloudFormation.
- Lead Kubernetes operations including cluster upgrades, capacity planning, and GitOps deployments.
- Evolve CI/CD pipelines across Jenkins, GitHub Actions, and ArgoCD to reduce manual steps.
- Drive observability stack enhancements using Prometheus, Grafana, and CloudWatch.
- Identify and execute FinOps initiatives to right-size infrastructure and reduce costs.
- Strengthen security posture through IAM, secret management, and regular audits.
- Participate in on-call rotations and lead post-incident analysis.
View Full Description & ApplyYou'll be redirected to the employer's site