Staff Site Reliability Engineer - Observability

New
J
JobgetherCloud Infrastructure
IndiaFull-TimeStaff
SalaryCompetitive compensation package with performance-based bonus and incentive opportunities.
Apply NowOpens the employer's application page

Job Details

Experience
8+ years of experience designing, operating, and supporting large-scale cloud infrastructure
Required Skills
AWSPythonElasticSearchKubernetesGoGrafanaPrometheusLinuxTerraform

Requirements

  • 8+ years of experience designing, operating, and supporting large-scale cloud infrastructure, with significant expertise in AWS environments.
  • 5+ years of hands-on experience with Kubernetes platforms such as EKS, AKS, GKE, Fargate, or similar container orchestration technologies.
  • 4+ years of Linux systems administration experience in production environments.
  • Strong programming skills in Go, Python, Ruby, or comparable languages used for automation and platform engineering.
  • Proven expertise with Infrastructure as Code technologies such as Terraform, Ansible, AWS CDK, or equivalent tools.
  • Strong understanding of observability platforms, distributed systems monitoring, logging, tracing, and performance optimization.
  • Experience with CI/CD pipelines, Git-based version control, automated testing, and modern software delivery practices.
  • Knowledge of SQL databases such as MySQL or PostgreSQL, networking fundamentals, distributed infrastructure, and cybersecurity best practices.
  • Excellent problem-solving, communication, collaboration, and mentoring skills.

Responsibilities

  • Design, implement, and continuously improve enterprise observability solutions using modern monitoring, logging, and tracing technologies, including Prometheus, Grafana, OpenTelemetry, Elastic, and related platforms.
  • Lead initiatives that improve the reliability, scalability, and performance of cloud-native and hybrid infrastructure across AWS and Kubernetes environments.
  • Manage vulnerability remediation, patch compliance, and infrastructure health across large-scale on-premises and cloud deployments while ensuring service-level objectives are achieved.
  • Operate, enhance, and optimize Kubernetes platforms and GitOps workflows, including ArgoCD and multi-tenant containerized environments.
  • Develop automation solutions and AI-assisted operational tooling that streamline incident response, reduce manual effort, and improve operational efficiency.
  • Partner with software engineering teams to establish monitoring standards, define meaningful service metrics, optimize alerting strategies, and improve production readiness.
  • Participate in incident management, root cause analysis, and continuous improvement initiatives while mentoring engineers and promoting Site Reliability Engineering best practices.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive compensation package with performance-based bonus and incentive opportunities.
Apply Now