Manager, Site Reliability Engineering
New
A
Aya HealthcareHealthcare Workforce Solutions
Remote, USFull-TimeManager
Salary$230,000 to $255,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years in a combination of Site Reliability Engineering, DevOps, Platform Engineering, or related production-operations roles.
- Required Skills
- AzureDevOpsSaaSDatadogHIPAA
Requirements
- 10+ years in a combination of Site Reliability Engineering, DevOps, Platform Engineering, or related production-operations roles.
- 4+ years of direct people management experience including hiring, performance management, and career development.
- Demonstrated ownership of reliability outcomes for customer-facing SaaS at scale.
- Deep Azure experience (3+ years) with hands-on depth in AKS, networking, identity, and platform services (or equivalent AWS/GCP depth).
- Modern observability fluency using Datadog, New Relic, Dynatrace, or AppDynamics.
- Hands-on experience integrating AI/LLM-assisted tooling into operational workflows.
- Proven experience as an incident commander leading severity-1 events and running blameless reviews.
- Experience operating with HIPAA, PHI, SOC 2, or comparable compliance constraints.
- Executive-grade communication skills for translating reliability work into business outcomes.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent.
Responsibilities
- Lead, mentor, and grow a team of high-performing Site Reliability Engineers across hiring, performance management, career development, and on-call rotation health.
- Set the operating cadence for the team including standups, incident reviews, SLO/error-budget reviews, post-incident learning, and capacity planning.
- Define reliability strategy for customer-facing products and internal platforms, operationalizing SLOs/SLIs and error budgets.
- Lead major incident response as senior incident commander and ensure systemic fixes for severity-1 events.
- Build the AIOps practice including anomaly detection, predictive alerting, and automated triage to reduce MTTD and MTTR.
- Operationalize AI-assisted workflows for runbook generation, log analysis, and change risk scoring.
- Partner with FinOps to drive platform unit economics, cost-to-serve, and capacity efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site