Staff Site Reliability Engineer
New
F
FilevineLegal AI
United StatesFull-TimeStaff
Salary235,000 - 275,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE, including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives for distributed production systems.
- Required Skills
- PythonKubernetesDistributed Systems
Requirements
- Bring 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE.
- Have 6+ years of experience in SRE.
- Have 3+ years leading complex, cross-functional technical initiatives for distributed production systems.
- Demonstrate expert-level depth in observability and platform infrastructure.
- Bring broad expertise in incident response, capacity planning, automation, and reliability engineering.
- Have advanced experience with a major container-orchestration platform, preferably Kubernetes.
- Have experience with an observability platform such as New Relic, Datadog, or equivalent.
- Demonstrate strong software-engineering ability in Python, Go, Bash, or another general-purpose language.
- Have experience building production tooling, automation, or platform capabilities.
- Be able to mentor engineers and communicate technical risk clearly to engineering, product, and executive audiences.
- Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is strongly preferred.
Responsibilities
- Define and execute technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
- Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
- Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.
- Lead the organization through complex production incidents and turn post-incident learning into permanent engineering improvements.
- Build self-service platform capabilities that reduce toil, improve engineering safety and velocity, and enable teams to own their reliability.
- Mentor engineers and guide long-term reliability and platform direction.
View Full Description & ApplyYou'll be redirected to the employer's site