Senior Site Reliability Engineer

New
J
JobgetherAI/ML Geospatial
CanadaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
PythonGCPKubernetesGoGrafanaPrometheusTerraform

Requirements

  • 3+ years of professional experience in Site Reliability Engineering, production engineering, or a closely related infrastructure role.
  • Strong hands-on experience with Google Cloud Platform, including cloud cost optimization, governance, and infrastructure management.
  • Proven experience managing Kubernetes clusters and workloads in production environments.
  • Experience with Infrastructure as Code tools such as Terraform or Google Cloud Deployment Manager.
  • Strong scripting and automation capabilities using Python, Bash, Go, or comparable languages.
  • Extensive experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, centralized logging, and distributed tracing.
  • Practical experience with incident management, post-incident analysis, production troubleshooting, and on-call operations.
  • Ability to define, implement, and monitor SLOs, SLIs, and error budgets.
  • Strong understanding of reliability engineering principles and a proactive approach to identifying and resolving operational risks.
  • Familiarity with DORA metrics and their application to engineering workflows is an asset.
  • Experience working in AI/ML, geospatial technology, or other data-intensive technical environments is considered a plus.
  • Strong communication and collaboration skills, with the ability to work effectively across distributed Product and Engineering teams.

Responsibilities

  • Design, build, and continuously evolve scalable and reliable cloud infrastructure on Google Cloud Platform.
  • Develop internal tooling, automation, and self-service capabilities that increase engineering efficiency and reduce operational dependencies.
  • Strengthen the observability platform across metrics, logging, distributed tracing, and monitoring to improve system visibility and reduce mean time to recovery.
  • Establish infrastructure cost visibility, governance, and optimization initiatives to improve cloud efficiency and manage spending responsibly.
  • Define and champion reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and DORA metrics.
  • Lead incident management activities, coordinate response efforts, facilitate post-incident reviews, and drive meaningful improvements based on incident learnings.
  • Participate in the on-call rotation and help ensure production systems remain stable, available, and resilient.
  • Partner with Product and Engineering teams to improve operational practices, deployment reliability, and overall software delivery performance.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now