Director, Site Reliability Engineering

New
J
JobgetherTechnology
CanadaFull-TimeDirector
Salary$243,800 USD
Apply NowOpens the employer's application page

Job Details

Experience
10+ years of relevant professional experience; 4+ years of experience leading SRE or comparable engineering teams
Required Skills
DockerPythonTypeScriptGoLinuxDistributed Systems

Requirements

  • 10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields.
  • 4+ years of experience leading SRE or comparable engineering teams.
  • Experience participating in or managing 24/7 on-call operations for large-scale production environments.
  • Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems.
  • Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments.
  • Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities.
  • Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions.
  • Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose.
  • Experience with high-level programming languages such as Go, Perl, TypeScript, Python, or comparable technologies.
  • Experience with AI-driven software development, including the design and implementation of agentic workflows.

Responsibilities

  • Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems.
  • Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices.
  • Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem.
  • Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation.
  • Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks.
  • Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions.
  • Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability.
View Full Description & ApplyYou'll be redirected to the employer's site
$243,800 USD
Apply Now