Expert Automation & Observability Engineer
New
E
EnsonoManaged Services
If you are not required to be on a client site, you can choose to work from home or in our Ensono officesFull-TimeSenior
Salary$140,000 to $180,000 annually
Apply NowOpens the employer's application page
Job Details
- Experience
- 12+ years of total IT experience, with a minimum of 5 to 7 years functioning as a Lead Architect, SRE, or Principal Observability Engineer
- Required Skills
- PythonKubernetesGrafanaPrometheusTerraformAnsibleServiceNow
Requirements
- 12+ years of total IT experience.
- 5 to 7 years functioning as a Lead Architect, SRE, or Principal Observability Engineer in a massive enterprise environment.
- Proven track record of migrating organizations from legacy monitoring to proactive, automated observability platforms.
- Hands-on expertise in building scalable, secure telemetry pipelines and time-series databases.
- Extensive experience leading FOAK rollouts and complex vendor/operations transition programs.
- Proficiency with IBM Instana, Grafana, Prometheus, OpenTelemetry, Telegraf, and InfluxDB.
- Experience with legacy monitoring tools like SolarWinds, Netcool, Elastic, or Splunk.
- Strong background in Kubernetes, Docker, OpenShift, and multi-cloud (AWS/Azure/GCP).
- Automation and DevOps skills including Ansible, Terraform, Python, Bash, and CI/CD tools.
- Understanding of ITIL 4 and advanced Major Incident Management.
- Preferred: CKA, Cloud Architect (AWS/Azure), or specific APM/Observability vendor certifications.
Responsibilities
- Architect and govern a unified observability framework covering metrics, logs, traces, and events using IBM Instana, Grafana, OpenTelemetry, Telegraf, and InfluxDB.
- Lead First-of-a-Kind (FOAK) implementations—evaluating new observability tech and converting them into secure, repeatable, production-ready patterns.
- Define and govern Service Level Indicators (SLIs), Objectives (SLOs), and error budgets.
- Serve as the senior technical escalation point, leading major P1/P2 incident war rooms and conducting evidence-based Root Cause Analysis (RCA).
- Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps.
- Design deep observability for Docker, Kubernetes, microservices, and multi-cloud environments.
- Lead complex Knowledge Transfer (KT) programs, vendor transitions, and operational readiness handovers.
View Full Description & ApplyYou'll be redirected to the employer's site