Site Reliability Engineer 3

New
J
JobgetherCloud Infrastructure
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
AWSGCPAzureLinuxTerraformAnsible

Requirements

  • 6+ years of experience in Site Reliability Engineering, AIOps, DevOps, or production engineering roles within large-scale cloud environments.
  • Strong knowledge of Linux/Unix systems, networking, distributed systems, and cloud platforms such as AWS, Azure, or Google Cloud.
  • Expert-level experience with ELK/OpenSearch, including log ingestion, indexing, scaling, querying, dashboards, and alerting.
  • Strong understanding of observability practices across logs, metrics, and distributed tracing.
  • Experience preparing telemetry data for AIOps implementation, including metadata enrichment, service mapping, normalization, and correlation.
  • Hands-on experience implementing AIOps solutions such as anomaly detection, alert correlation, intelligent alerting, and automated incident enrichment.
  • Experience integrating AI capabilities into SRE workflows, including LLM-based incident analysis, operational assistants, or automated troubleshooting processes.
  • Understanding of prompt engineering concepts for operational use cases.
  • Strong knowledge of incident management, root-cause analysis, SLOs, reliability practices, and production operations.
  • Experience with Infrastructure as Code tools such as Terraform or Ansible.

Responsibilities

  • Provide end-to-end reliability ownership for production systems, including on-call support, incident response, postmortems, and SLO/SLI management.
  • Build and maintain observability solutions across metrics, logs, and traces with intelligent monitoring and anomaly detection capabilities.
  • Design and implement AIOps workflows for event ingestion, enrichment, correlation, deduplication, and alert noise reduction.
  • Develop AI-assisted operational processes, including automated log analysis, incident summaries, root-cause analysis support, and runbook generation.
  • Create automation solutions that reduce operational effort through AI-enhanced runbooks, controlled remediation workflows, and self-healing capabilities.
  • Implement safeguards, approval processes, rollback strategies, and audit controls for automated operational actions.
  • Own and improve observability platforms, including ELK/OpenSearch environments for logging, indexing, querying, dashboards, and alerting.
  • Enhance incident response processes through AI-assisted diagnosis, timeline reconstruction, and operational intelligence.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now