Senior Site Reliability Engineer
New
S
ShopmonkeyAutomotive Software
Remote within Europe, Must consistently overlap with US Eastern Time morningsContractSenior
Salary53 USD per hour
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 5+ years
- Required Skills
- Node.jsGCPKubernetesGoGrafanaPrometheusTerraformHelm
Requirements
- 5+ years of professional experience in Site Reliability Engineering, Platform Engineering, or Observability.
- Proven hands-on experience building an OpenTelemetry pipeline in a production GCP environment.
- Strong experience with Google Kubernetes Engine (GKE).
- Experience with Google Cloud Monitoring, Cloud Trace, and Managed Service for Prometheus.
- Strong knowledge of Grafana, Prometheus, distributed tracing, metrics, logging, and telemetry architecture.
- Practical experience defining SLIs/SLOs and implementing multi-window, multi-burn-rate alerting.
- Strong hands-on experience managing infrastructure with Terraform and Helm.
- Production-level Kubernetes experience.
- Experience participating in on-call rotations, responding to incidents, and conducting root-cause analyses.
- Ability to read, troubleshoot, and submit production code changes in Go or Node.js.
- Strong written and verbal English communication skills.
Responsibilities
- Design, build, and improve OpenTelemetry pipelines for metrics, traces, and logs.
- Strengthen observability across services running on GCP and GKE.
- Integrate telemetry with tools such as Google Cloud Monitoring, Cloud Trace, Managed Service for Prometheus, and Grafana.
- Define and implement SLIs and SLOs based on customer and service reliability goals.
- Build actionable burn-rate alerts to identify reliability risks.
- Manage observability infrastructure through Terraform, Helm, and Kubernetes.
- Participate in on-call and incident-response activities.
- Contribute to root-cause analyses and implement preventative improvements.
- Collaborate with software engineers and open pull requests against Go or Node.js services.
- Document operational standards and promote observability practices.
View Full Description & ApplyYou'll be redirected to the employer's site