Senior Site Reliability Engineer
New
J
JobgetherAI/ML Geospatial
CanadaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- PythonGCPKubernetesGoGrafanaPrometheusTerraform
Requirements
- 3+ years of professional experience in Site Reliability Engineering, production engineering, or a closely related infrastructure role.
- Strong hands-on experience with Google Cloud Platform, including cloud cost optimization, governance, and infrastructure management.
- Proven experience managing Kubernetes clusters and workloads in production environments.
- Experience with Infrastructure as Code tools such as Terraform or Google Cloud Deployment Manager.
- Strong scripting and automation capabilities using Python, Bash, Go, or comparable languages.
- Extensive experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, centralized logging, and distributed tracing.
- Practical experience with incident management, post-incident analysis, production troubleshooting, and on-call operations.
- Ability to define, implement, and monitor SLOs, SLIs, and error budgets.
- Strong understanding of reliability engineering principles and a proactive approach to identifying and resolving operational risks.
- Familiarity with DORA metrics and their application to engineering workflows is an asset.
- Experience working in AI/ML, geospatial technology, or other data-intensive technical environments is considered a plus.
- Strong communication and collaboration skills, with the ability to work effectively across distributed Product and Engineering teams.
Responsibilities
- Design, build, and continuously evolve scalable and reliable cloud infrastructure on Google Cloud Platform.
- Develop internal tooling, automation, and self-service capabilities that increase engineering efficiency and reduce operational dependencies.
- Strengthen the observability platform across metrics, logging, distributed tracing, and monitoring to improve system visibility and reduce mean time to recovery.
- Establish infrastructure cost visibility, governance, and optimization initiatives to improve cloud efficiency and manage spending responsibly.
- Define and champion reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and DORA metrics.
- Lead incident management activities, coordinate response efforts, facilitate post-incident reviews, and drive meaningful improvements based on incident learnings.
- Participate in the on-call rotation and help ensure production systems remain stable, available, and resilient.
- Partner with Product and Engineering teams to improve operational practices, deployment reliability, and overall software delivery performance.
View Full Description & ApplyYou'll be redirected to the employer's site