Senior Software Engineer, Site Reliability Engineering

New
J
JobgetherSite reliability engineering
Remote work opportunity from eligible locations in Ontario and British Columbia, Canada.Full-TimeSenior
SalaryExpected total cash compensation of CAD $180,200–$233,200
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience managing infrastructure and systems
Required Skills
AWSLinuxMicroservicesDistributed Systems

Requirements

  • Bring 5+ years of experience managing infrastructure and systems, ideally in large-scale or distributed production environments.
  • Have extensive hands-on expertise with AWS and Linux-based systems.
  • Be able to read, write, debug, and maintain production-facing software and systems.
  • Understand large-scale distributed systems and web technologies, including DNS, TLS, HTTP/S, and TCP/IP.
  • Have experience operating, instrumenting, and observing distributed microservices in production cloud environments.
  • Have experience designing resilient, scalable infrastructure and platform services focused on availability, performance, and reliability.
  • Be able to break down complex technical challenges and make thoughtful trade-offs based on business and engineering impact.
  • Work effectively with technical and non-technical stakeholders across organizational levels.
  • Operate independently while contributing to a collaborative, cross-functional engineering environment.
  • Take ownership of production systems, operational excellence, and continuous improvement.

Responsibilities

  • Design, develop, and maintain software and infrastructure to improve service availability, scalability, performance, and operational efficiency.
  • Establish architectural direction for infrastructure and platform services and provide technical guidance to engineering teams.
  • Build and improve tools, processes, and systems for deployment, infrastructure, service, and change management.
  • Troubleshoot and resolve complex production issues across the software development lifecycle.
  • Develop platform capabilities that help engineering teams build, deploy, operate, and observe services.
  • Conduct capacity planning and demand forecasting to identify growth needs and performance bottlenecks.
  • Instrument, operate, and monitor distributed microservices and cloud-based systems.
  • Participate in a rotating on-call schedule and contribute to incident response, service recovery, and reliability improvements.
  • Partner with cross-functional engineering teams and stakeholders to identify opportunities and deliver platform solutions.
View Full Description & ApplyYou'll be redirected to the employer's site
Expected total cash compensation of CAD $180,200–$233,200
Apply Now