Senior Software Engineer, Site Reliability Engineering
New
J
JobgetherSite reliability engineering
Remote work opportunity from eligible locations in Ontario and British Columbia, Canada.Full-TimeSenior
SalaryExpected total cash compensation of CAD $180,200–$233,200
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience managing infrastructure and systems
- Required Skills
- AWSLinuxMicroservicesDistributed Systems
Requirements
- Bring 5+ years of experience managing infrastructure and systems, ideally in large-scale or distributed production environments.
- Have extensive hands-on expertise with AWS and Linux-based systems.
- Be able to read, write, debug, and maintain production-facing software and systems.
- Understand large-scale distributed systems and web technologies, including DNS, TLS, HTTP/S, and TCP/IP.
- Have experience operating, instrumenting, and observing distributed microservices in production cloud environments.
- Have experience designing resilient, scalable infrastructure and platform services focused on availability, performance, and reliability.
- Be able to break down complex technical challenges and make thoughtful trade-offs based on business and engineering impact.
- Work effectively with technical and non-technical stakeholders across organizational levels.
- Operate independently while contributing to a collaborative, cross-functional engineering environment.
- Take ownership of production systems, operational excellence, and continuous improvement.
Responsibilities
- Design, develop, and maintain software and infrastructure to improve service availability, scalability, performance, and operational efficiency.
- Establish architectural direction for infrastructure and platform services and provide technical guidance to engineering teams.
- Build and improve tools, processes, and systems for deployment, infrastructure, service, and change management.
- Troubleshoot and resolve complex production issues across the software development lifecycle.
- Develop platform capabilities that help engineering teams build, deploy, operate, and observe services.
- Conduct capacity planning and demand forecasting to identify growth needs and performance bottlenecks.
- Instrument, operate, and monitor distributed microservices and cloud-based systems.
- Participate in a rotating on-call schedule and contribute to incident response, service recovery, and reliability improvements.
- Partner with cross-functional engineering teams and stakeholders to identify opportunities and deliver platform solutions.
View Full Description & ApplyYou'll be redirected to the employer's site