Director, Site Reliability Engineering
New
J
JobgetherTechnology
CanadaFull-TimeDirector
Salary$243,800 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of relevant professional experience; 4+ years of experience leading SRE or comparable engineering teams
- Required Skills
- DockerPythonTypeScriptGoLinuxDistributed Systems
Requirements
- 10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields.
- 4+ years of experience leading SRE or comparable engineering teams.
- Experience participating in or managing 24/7 on-call operations for large-scale production environments.
- Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems.
- Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments.
- Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities.
- Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions.
- Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose.
- Experience with high-level programming languages such as Go, Perl, TypeScript, Python, or comparable technologies.
- Experience with AI-driven software development, including the design and implementation of agentic workflows.
Responsibilities
- Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems.
- Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices.
- Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem.
- Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation.
- Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks.
- Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions.
- Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability.
View Full Description & ApplyYou'll be redirected to the employer's site