Site Reliability Engineer 3
New
J
JobgetherCloud Infrastructure
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- AWSGCPAzureLinuxTerraformAnsible
Requirements
- 6+ years of experience in Site Reliability Engineering, AIOps, DevOps, or production engineering roles within large-scale cloud environments.
- Strong knowledge of Linux/Unix systems, networking, distributed systems, and cloud platforms such as AWS, Azure, or Google Cloud.
- Expert-level experience with ELK/OpenSearch, including log ingestion, indexing, scaling, querying, dashboards, and alerting.
- Strong understanding of observability practices across logs, metrics, and distributed tracing.
- Experience preparing telemetry data for AIOps implementation, including metadata enrichment, service mapping, normalization, and correlation.
- Hands-on experience implementing AIOps solutions such as anomaly detection, alert correlation, intelligent alerting, and automated incident enrichment.
- Experience integrating AI capabilities into SRE workflows, including LLM-based incident analysis, operational assistants, or automated troubleshooting processes.
- Understanding of prompt engineering concepts for operational use cases.
- Strong knowledge of incident management, root-cause analysis, SLOs, reliability practices, and production operations.
- Experience with Infrastructure as Code tools such as Terraform or Ansible.
Responsibilities
- Provide end-to-end reliability ownership for production systems, including on-call support, incident response, postmortems, and SLO/SLI management.
- Build and maintain observability solutions across metrics, logs, and traces with intelligent monitoring and anomaly detection capabilities.
- Design and implement AIOps workflows for event ingestion, enrichment, correlation, deduplication, and alert noise reduction.
- Develop AI-assisted operational processes, including automated log analysis, incident summaries, root-cause analysis support, and runbook generation.
- Create automation solutions that reduce operational effort through AI-enhanced runbooks, controlled remediation workflows, and self-healing capabilities.
- Implement safeguards, approval processes, rollback strategies, and audit controls for automated operational actions.
- Own and improve observability platforms, including ELK/OpenSearch environments for logging, indexing, querying, dashboards, and alerting.
- Enhance incident response processes through AI-assisted diagnosis, timeline reconstruction, and operational intelligence.
View Full Description & ApplyYou'll be redirected to the employer's site