Staff Site Reliability Engineer

J
JobgetherCloud Infrastructure
Based in the United StatesFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
12+ years of experience in software engineering, infrastructure engineering, platform engineering, or Site Reliability Engineering; 6 years of dedicated SRE experience; 3+ years leading complex, cross-functional technical initiatives
Required Skills
PythonBashKubernetesGoDatadogDistributed Systems

Requirements

  • 12+ years of experience in software engineering, infrastructure engineering, platform engineering, or Site Reliability Engineering.
  • At least 6 years of dedicated SRE experience.
  • 3+ years leading complex, cross-functional technical initiatives involving large-scale distributed systems.
  • Deep expertise in observability, infrastructure reliability, incident response, capacity management, automation, and operational excellence.
  • Advanced experience with container orchestration technologies, particularly Kubernetes.
  • Experience with modern observability platforms such as Datadog, New Relic, or similar solutions.
  • Strong software engineering skills with proficiency in languages such as Python, Go, Bash, or other general-purpose programming languages.
  • Proven experience building automation frameworks, internal tooling, platform services, or Infrastructure-as-Code solutions.
  • Solid understanding of cloud-native architectures, distributed systems design, and production-scale operational challenges.
  • Demonstrated ability to mentor engineers, influence technical direction, and communicate effectively with technical and non-technical stakeholders, including executive leadership.
  • Experience applying AI/ML concepts to operational workflows, observability, or platform management is highly desirable.
  • Familiarity with compliance-focused or regulated environments such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is preferred.

Responsibilities

  • Define and execute the technical strategy for observability, alerting, platform infrastructure, and overall operational excellence.
  • Lead the design and evolution of scalable, secure, reliable, and cost-efficient cloud-native platforms and distributed systems.
  • Establish and champion reliability best practices, including SLIs, SLOs, error budgets, capacity planning, operational readiness reviews, and automation initiatives.
  • Guide teams through complex production incidents, driving effective response processes and ensuring lessons learned translate into lasting engineering improvements.
  • Build and promote self-service platform capabilities that reduce operational toil, improve developer productivity, and strengthen system reliability.
  • Drive adoption of AI and machine learning technologies for observability, anomaly detection, incident management, automated remediation, and resource optimization.
  • Partner with engineering leadership on critical architectural decisions and long-term platform strategy.
  • Mentor engineers across multiple levels, providing technical guidance and fostering a culture of reliability ownership and continuous improvement.
  • Serve as a trusted advisor on production risk, operational health, and infrastructure scalability across the organization.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now