Senior Site Reliability Engineer
New
C
CamundaSoftware Development
Fully remote and globalFull-TimeSenior
SalaryUnited States: $149,800.00 to $241,500.00
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSPythonGCPKubernetesGoGrafanaPrometheusTerraform
Requirements
- Deep hands-on experience with Kubernetes, including managing workloads, networking, and storage at scale.
- Infrastructure as code expertise using tools like Terraform.
- Proven experience in monitoring and observability using tools like Prometheus or Grafana.
- Strong 3rd-level support and incident response skills with ability to conduct root cause analysis.
- A passion for automation and building reliable, maintainable systems.
- Demonstrated ability to use AI tools for research, coding, and documentation with human oversight.
- Experience with major cloud providers such as AWS (EKS) or GCP (GKE) is preferred.
- Familiarity with ArgoCD or GitOps workflows is preferred.
- Proficiency in Python, Go, or similar languages for scripting and automation is preferred.
- Experience with defining SLOs and configuring alerting frameworks is preferred.
Responsibilities
- Evolve and maintain the Kubernetes-based, multi-cloud platform architecture to ensure availability, scalability, and fault tolerance.
- Implement and improve monitoring and alerting tools to provide system health visibility to SREs and developers.
- Participate in on-call rotations and manage end-to-end system ownership, including creating runbooks and automation.
- Work cross-functionally with product engineering, management, and support to deliver features.
- Identify and automate repetitive manual work to improve system quality.
- Mentor less experienced engineers on complex infrastructure challenges.
View Full Description & ApplyYou'll be redirected to the employer's site