Senior Site Reliability Engineer

New
J
JobgetherAI/ML Geospatial
United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
PythonBashGCPKubernetesGoGrafanaPrometheusTerraform

Requirements

  • 3+ years of professional experience in Site Reliability Engineering, production engineering, or a closely related infrastructure role.
  • Strong hands-on experience with Google Cloud Platform, including cloud cost optimization, governance, and infrastructure management.
  • Proven experience managing Kubernetes clusters and workloads in production environments.
  • Experience with Infrastructure as Code tools such as Terraform or Google Cloud Deployment Manager.
  • Strong scripting and automation capabilities using Python, Bash, Go, or comparable languages.
  • Extensive experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, centralized logging, and distributed tracing.
  • Practical experience with incident management, post-incident analysis, production troubleshooting, and on-call operations.
  • Ability to define, implement, and monitor SLOs, SLIs, and error budgets.
  • Strong understanding of reliability engineering principles and a proactive approach to identifying and resolving operational risks.

Responsibilities

  • Design, build, and continuously evolve scalable and reliable cloud infrastructure on Google Cloud Platform.
  • Develop internal tooling, automation, and self-service capabilities that increase engineering efficiency and reduce operational dependencies.
  • Strengthen the observability platform across metrics, logging, distributed tracing, and monitoring to improve system visibility and reduce mean time to recovery.
  • Establish infrastructure cost visibility, governance, and optimization initiatives to improve cloud efficiency and manage spending responsibly.
  • Define and champion reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and DORA metrics.
  • Lead incident management activities, coordinate response efforts, facilitate post-incident reviews, and drive meaningful improvements based on incident learnings.
  • Participate in the on-call rotation and help ensure production systems remain stable, available, and resilient.
  • Partner with Product and Engineering teams to improve operational practices, deployment reliability, and overall software delivery performance.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now