Staff Site Reliability Engineer-Observability
New
G
GoDaddyCloud Infrastructure
India, RemoteFull-TimeStaff
Salary6,370,000 - 11,830,000 INR per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years of hands-on experience working with AWS
- Required Skills
- AWSPythonKubernetesRubyGoGrafanaPrometheusLinuxTerraformAnsible
Requirements
- 8+ years of hands-on experience working with AWS, in cloud-native and agnostic capacities.
- 5+ years of experience with Kubernetes (EKS/AKS/GKE/Fargate, etc.).
- 4+ years of expertise in Linux administration.
- Strong coding skills in languages such as Go, Python, Ruby.
- Strong experience in coding infrastructure as code (Terraform, Ansible, CDK, etc.).
- Understanding of CI/CD concepts, version control systems, and testing tools (e.g., Jenkins, Gradle, Maven).
- Hands-on experience with SQL databases, especially MySQL or PostgreSQL.
- Verified background in developing and maintaining distributed infrastructure and systems.
- Understanding of networking principles and dedication to cybersecurity guidelines.
- Ability to work independently and perform effectively under pressure.
Responsibilities
- Build and evolve observability using Prometheus/Mimir, the Grafana LGTM stack, Elastic/OTEL, and Site24x7.
- Steer the cloud migration journey and maintain service reliability and performance.
- Own patching compliance and vulnerability remediation across a mixed on-prem and AWS fleet.
- Operate and extend the multi-tenant Kubernetes/ArgoCD platform.
- Contribute to AI/automation initiatives including auto-generating runbooks and ServiceNow change-risk scoring.
- Consult with partner dev teams on metrics, alert thresholds, and monitoring standards.
- Mentor other engineers and mature operational practices.
View Full Description & ApplyYou'll be redirected to the employer's site