Staff Site Reliability Engineer
New
F
FingerprintFraud Detection
Fingerprint is a globally dispersed, 100% remote company. We are an all-remote company and people can join our workforce from almost any country.Full-TimeStaff
Salary177,000 - 240,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of engineering experience, with 3+ years as an SRE, production engineer, or reliability-focused Staff engineer
- Required Skills
- AWSKubernetesTypeScriptGoDatadogDistributed Systems
Requirements
- 10+ years of engineering experience.
- 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams.
- Proven experience with SLI/SLO design and error budget implementation.
- Strong incident leadership skills in high-severity, customer-facing environments.
- Hands-on depth in distributed systems failure modes (cache/database saturation, cascading failure, retry storms).
- Fluent in Kubernetes, AWS, and modern observability tooling (e.g., Datadog).
- Proficiency in reading/writing production code (Go, TypeScript, or similar) and infrastructure as code.
- Demonstrated ability to lead through influence across non-managed teams.
- Experience coaching engineers to own reliability outcomes.
- Exceptional written communication skills for async, documented decision-making.
- Experience using AI tools for incident investigation, analysis, and building operational tooling.
Responsibilities
- Define SLIs and SLOs for critical request paths (identification, events, server APIs, client agents) and tie them to team decisions.
- Introduce error budgets to balance reliability investments against feature work.
- Strengthen incident lifecycles, from detection and communication to postmortem quality and follow-up.
- Lead reliability reviews for high-risk changes and implement failure testing (game days, chaos exercises).
- Develop Staff and Lead engineers as reliability leaders within their groups.
- Codify lightweight production readiness, on-call standards, and runbook practices.
- Lead AI adoption for incident investigation, observability, and runbook automation.
- Dig into production during incidents and build tooling, dashboards, or reference implementations.
View Full Description & ApplyYou'll be redirected to the employer's site