Senior Observability Engineer / Platform Engineer

New
J
JobgetherCloud Engineering
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
AWSPythonGCPKubernetesAzureGrafanaPrometheusTerraform

Requirements

  • 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, Observability Engineering, or a closely related discipline.
  • Strong hands-on expertise with observability technologies such as Prometheus, Grafana, Splunk, OpenSearch, Elastic, or equivalent platforms.
  • Experience with distributed tracing solutions such as Jaeger, Tempo, and OpenTelemetry, with a solid understanding of modern observability practices.
  • Strong knowledge of Kubernetes, Docker, and containerized workloads, ideally gained through enterprise-scale production environments.
  • Practical experience working with AWS, Azure, or GCP cloud environments.
  • Demonstrated ability to manage incidents, tune alerts, troubleshoot complex production issues, and improve operational reliability.
  • Strong scripting and automation capabilities using Python, Shell, or similar languages.
  • Familiarity with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or ArgoCD.
  • Understanding of SRE practices, including SLO/SLI frameworks and error budgets, is preferred.
  • Exposure to AIOps, intelligent incident response, automated remediation, and OpenTelemetry implementations is an advantage.
  • Experience working with enterprise-scale production platforms and an awareness of security, compliance, and governance considerations in cloud-native environments.
  • Strong collaboration, communication, analytical, and problem-solving skills, with the ability to work effectively across engineering and operations teams.

Responsibilities

  • Design, implement, and maintain enterprise observability platforms covering metrics, logs, traces, and events.
  • Build and manage observability solutions using platforms such as Prometheus, Grafana, OpenSearch, Splunk, Elastic, Datadog, Dynatrace, New Relic, or equivalent technologies.
  • Develop dashboards, SLOs, SLIs, alerting rules, and service health monitoring frameworks to provide actionable visibility into production environments.
  • Integrate monitoring and observability capabilities across Kubernetes, containerized workloads, and public cloud platforms.
  • Enable effective incident management, root cause analysis, and production troubleshooting through robust observability practices.
  • Automate monitoring configuration, platform onboarding, and operational processes using Infrastructure-as-Code and CI/CD pipelines.
  • Partner with engineering, platform, SRE, DevOps, and operations teams to improve system reliability, performance, scalability, and availability.
  • Contribute to observability maturity initiatives, including distributed tracing, OpenTelemetry, AIOps, intelligent alerting, and automated remediation.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now