Manager, Site Reliability Engineering

New
A
Aya HealthcareHealthcare Workforce Solutions
Remote, USFull-TimeManager
Salary$230,000 to $255,000
Apply NowOpens the employer's application page

Job Details

Experience
10+ years in a combination of Site Reliability Engineering, DevOps, Platform Engineering, or related production-operations roles.
Required Skills
AzureDevOpsSaaSDatadogHIPAA

Requirements

  • 10+ years in a combination of Site Reliability Engineering, DevOps, Platform Engineering, or related production-operations roles.
  • 4+ years of direct people management experience including hiring, performance management, and career development.
  • Demonstrated ownership of reliability outcomes for customer-facing SaaS at scale.
  • Deep Azure experience (3+ years) with hands-on depth in AKS, networking, identity, and platform services (or equivalent AWS/GCP depth).
  • Modern observability fluency using Datadog, New Relic, Dynatrace, or AppDynamics.
  • Hands-on experience integrating AI/LLM-assisted tooling into operational workflows.
  • Proven experience as an incident commander leading severity-1 events and running blameless reviews.
  • Experience operating with HIPAA, PHI, SOC 2, or comparable compliance constraints.
  • Executive-grade communication skills for translating reliability work into business outcomes.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent.

Responsibilities

  • Lead, mentor, and grow a team of high-performing Site Reliability Engineers across hiring, performance management, career development, and on-call rotation health.
  • Set the operating cadence for the team including standups, incident reviews, SLO/error-budget reviews, post-incident learning, and capacity planning.
  • Define reliability strategy for customer-facing products and internal platforms, operationalizing SLOs/SLIs and error budgets.
  • Lead major incident response as senior incident commander and ensure systemic fixes for severity-1 events.
  • Build the AIOps practice including anomaly detection, predictive alerting, and automated triage to reduce MTTD and MTTR.
  • Operationalize AI-assisted workflows for runbook generation, log analysis, and change risk scoring.
  • Partner with FinOps to drive platform unit economics, cost-to-serve, and capacity efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site
$230,000 to $255,000
Apply Now