Staff Site Reliability Engineer, Ads
R
Reddit, Inc.Advertising Technology
This role is remote friendly.Full-TimeStaff
Salary$217,000 — $303,900 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- KafkaKubernetesGoBigQueryDistributed Systems
Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles.
- Proven experience operating large-scale distributed systems and high-traffic production environments.
- Strong software engineering skills in backend languages such as Go.
- Deep expertise in observability systems (metrics, logging, tracing, and alerting).
- Strong background in cloud-native architectures and scale engineering.
- Experience designing highly-available systems with strong operational and reliability practices.
- Demonstrated ability to drive operational excellence through SLOs, automation, and incident management.
- Experience troubleshooting complex issues across modern distributed system stacks.
- Strong cross-functional collaboration and communication skills to influence technical direction.
Responsibilities
- Lead reliability initiatives across ad serving, auctions, targeting, reporting, measurement, and billing domains.
- Partner with engineering leadership to develop roadmaps for reliability, scalability, and operational excellence.
- Design and build platforms, tooling, and automation to enhance system reliability and developer productivity.
- Drive architecture reviews and influence technical decisions for revenue-generating systems.
- Participate in on-call rotations and lead complex cross-functional incident response and investigations.
- Identify systemic reliability risks and implement long-term resilience solutions.
- Establish reliability metrics around advertiser-critical user journeys.
- Mentor engineers and provide technical leadership across multiple teams.
View Full Description & ApplyYou'll be redirected to the employer's site