Staff Site Reliability Engineer
J
JobgetherCloud Infrastructure
Based in the United StatesFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 12+ years of experience in software engineering, infrastructure engineering, platform engineering, or Site Reliability Engineering; 6 years of dedicated SRE experience; 3+ years leading complex, cross-functional technical initiatives
- Required Skills
- PythonBashKubernetesGoDatadogDistributed Systems
Requirements
- 12+ years of experience in software engineering, infrastructure engineering, platform engineering, or Site Reliability Engineering.
- At least 6 years of dedicated SRE experience.
- 3+ years leading complex, cross-functional technical initiatives involving large-scale distributed systems.
- Deep expertise in observability, infrastructure reliability, incident response, capacity management, automation, and operational excellence.
- Advanced experience with container orchestration technologies, particularly Kubernetes.
- Experience with modern observability platforms such as Datadog, New Relic, or similar solutions.
- Strong software engineering skills with proficiency in languages such as Python, Go, Bash, or other general-purpose programming languages.
- Proven experience building automation frameworks, internal tooling, platform services, or Infrastructure-as-Code solutions.
- Solid understanding of cloud-native architectures, distributed systems design, and production-scale operational challenges.
- Demonstrated ability to mentor engineers, influence technical direction, and communicate effectively with technical and non-technical stakeholders, including executive leadership.
- Experience applying AI/ML concepts to operational workflows, observability, or platform management is highly desirable.
- Familiarity with compliance-focused or regulated environments such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is preferred.
Responsibilities
- Define and execute the technical strategy for observability, alerting, platform infrastructure, and overall operational excellence.
- Lead the design and evolution of scalable, secure, reliable, and cost-efficient cloud-native platforms and distributed systems.
- Establish and champion reliability best practices, including SLIs, SLOs, error budgets, capacity planning, operational readiness reviews, and automation initiatives.
- Guide teams through complex production incidents, driving effective response processes and ensuring lessons learned translate into lasting engineering improvements.
- Build and promote self-service platform capabilities that reduce operational toil, improve developer productivity, and strengthen system reliability.
- Drive adoption of AI and machine learning technologies for observability, anomaly detection, incident management, automated remediation, and resource optimization.
- Partner with engineering leadership on critical architectural decisions and long-term platform strategy.
- Mentor engineers across multiple levels, providing technical guidance and fostering a culture of reliability ownership and continuous improvement.
- Serve as a trusted advisor on production risk, operational health, and infrastructure scalability across the organization.
View Full Description & ApplyYou'll be redirected to the employer's site