Senior Site Reliability Engineer

New
C
CamundaSoftware Development
Fully remote and globalFull-TimeSenior
SalaryUnited States: $149,800.00 to $241,500.00
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSPythonGCPKubernetesGoGrafanaPrometheusTerraform

Requirements

  • Deep hands-on experience with Kubernetes, including managing workloads, networking, and storage at scale.
  • Infrastructure as code expertise using tools like Terraform.
  • Proven experience in monitoring and observability using tools like Prometheus or Grafana.
  • Strong 3rd-level support and incident response skills with ability to conduct root cause analysis.
  • A passion for automation and building reliable, maintainable systems.
  • Demonstrated ability to use AI tools for research, coding, and documentation with human oversight.
  • Experience with major cloud providers such as AWS (EKS) or GCP (GKE) is preferred.
  • Familiarity with ArgoCD or GitOps workflows is preferred.
  • Proficiency in Python, Go, or similar languages for scripting and automation is preferred.
  • Experience with defining SLOs and configuring alerting frameworks is preferred.

Responsibilities

  • Evolve and maintain the Kubernetes-based, multi-cloud platform architecture to ensure availability, scalability, and fault tolerance.
  • Implement and improve monitoring and alerting tools to provide system health visibility to SREs and developers.
  • Participate in on-call rotations and manage end-to-end system ownership, including creating runbooks and automation.
  • Work cross-functionally with product engineering, management, and support to deliver features.
  • Identify and automate repetitive manual work to improve system quality.
  • Mentor less experienced engineers on complex infrastructure challenges.
View Full Description & ApplyYou'll be redirected to the employer's site
United States: $149,800.00 to $241,500.00
Apply Now