Senior Observability Engineer / Platform Engineer
New
J
JobgetherCloud Engineering
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- AWSPythonGCPKubernetesAzureGrafanaPrometheusTerraform
Requirements
- 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, Observability Engineering, or a closely related discipline.
- Strong hands-on expertise with observability technologies such as Prometheus, Grafana, Splunk, OpenSearch, Elastic, or equivalent platforms.
- Experience with distributed tracing solutions such as Jaeger, Tempo, and OpenTelemetry, with a solid understanding of modern observability practices.
- Strong knowledge of Kubernetes, Docker, and containerized workloads, ideally gained through enterprise-scale production environments.
- Practical experience working with AWS, Azure, or GCP cloud environments.
- Demonstrated ability to manage incidents, tune alerts, troubleshoot complex production issues, and improve operational reliability.
- Strong scripting and automation capabilities using Python, Shell, or similar languages.
- Familiarity with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or ArgoCD.
- Understanding of SRE practices, including SLO/SLI frameworks and error budgets, is preferred.
- Exposure to AIOps, intelligent incident response, automated remediation, and OpenTelemetry implementations is an advantage.
- Experience working with enterprise-scale production platforms and an awareness of security, compliance, and governance considerations in cloud-native environments.
- Strong collaboration, communication, analytical, and problem-solving skills, with the ability to work effectively across engineering and operations teams.
Responsibilities
- Design, implement, and maintain enterprise observability platforms covering metrics, logs, traces, and events.
- Build and manage observability solutions using platforms such as Prometheus, Grafana, OpenSearch, Splunk, Elastic, Datadog, Dynatrace, New Relic, or equivalent technologies.
- Develop dashboards, SLOs, SLIs, alerting rules, and service health monitoring frameworks to provide actionable visibility into production environments.
- Integrate monitoring and observability capabilities across Kubernetes, containerized workloads, and public cloud platforms.
- Enable effective incident management, root cause analysis, and production troubleshooting through robust observability practices.
- Automate monitoring configuration, platform onboarding, and operational processes using Infrastructure-as-Code and CI/CD pipelines.
- Partner with engineering, platform, SRE, DevOps, and operations teams to improve system reliability, performance, scalability, and availability.
- Contribute to observability maturity initiatives, including distributed tracing, OpenTelemetry, AIOps, intelligent alerting, and automated remediation.
View Full Description & ApplyYou'll be redirected to the employer's site