Staff Site Reliability Engineer - Observability
New
J
JobgetherCloud Infrastructure
IndiaFull-TimeStaff
SalaryCompetitive compensation package with performance-based bonus and incentive opportunities.
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years of experience designing, operating, and supporting large-scale cloud infrastructure
- Required Skills
- AWSPythonElasticSearchKubernetesGoGrafanaPrometheusLinuxTerraform
Requirements
- 8+ years of experience designing, operating, and supporting large-scale cloud infrastructure, with significant expertise in AWS environments.
- 5+ years of hands-on experience with Kubernetes platforms such as EKS, AKS, GKE, Fargate, or similar container orchestration technologies.
- 4+ years of Linux systems administration experience in production environments.
- Strong programming skills in Go, Python, Ruby, or comparable languages used for automation and platform engineering.
- Proven expertise with Infrastructure as Code technologies such as Terraform, Ansible, AWS CDK, or equivalent tools.
- Strong understanding of observability platforms, distributed systems monitoring, logging, tracing, and performance optimization.
- Experience with CI/CD pipelines, Git-based version control, automated testing, and modern software delivery practices.
- Knowledge of SQL databases such as MySQL or PostgreSQL, networking fundamentals, distributed infrastructure, and cybersecurity best practices.
- Excellent problem-solving, communication, collaboration, and mentoring skills.
Responsibilities
- Design, implement, and continuously improve enterprise observability solutions using modern monitoring, logging, and tracing technologies, including Prometheus, Grafana, OpenTelemetry, Elastic, and related platforms.
- Lead initiatives that improve the reliability, scalability, and performance of cloud-native and hybrid infrastructure across AWS and Kubernetes environments.
- Manage vulnerability remediation, patch compliance, and infrastructure health across large-scale on-premises and cloud deployments while ensuring service-level objectives are achieved.
- Operate, enhance, and optimize Kubernetes platforms and GitOps workflows, including ArgoCD and multi-tenant containerized environments.
- Develop automation solutions and AI-assisted operational tooling that streamline incident response, reduce manual effort, and improve operational efficiency.
- Partner with software engineering teams to establish monitoring standards, define meaningful service metrics, optimize alerting strategies, and improve production readiness.
- Participate in incident management, root cause analysis, and continuous improvement initiatives while mentoring engineers and promoting Site Reliability Engineering best practices.
View Full Description & ApplyYou'll be redirected to the employer's site