Staff Site Reliability Engineer, Ads
R
RedditAdvertising Technology
This role is remote friendly.Full-TimeStaff
Salary$217,000 — $303,900 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles
- Required Skills
- GoDistributed Systems
Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems.
- Strong experience evolving high traffic, user-facing production environments.
- Strong cross-functional collaborations skills to lead and influence projects driving operational excellence.
- Deep expertise in modern distributed systems, scale engineering, and cloud-native architectures.
- Experience designing highly-available systems with strong operational and reliability practices.
- Strong software engineering skills in languages like Go.
- Strong understanding of observability systems including metrics, logging, tracing, and alerting.
- Experience improving reliability through SLOs, automation, incident management, and performance optimization.
- Demonstrated ability to troubleshoot complex issues across a modern distributed system stack.
Responsibilities
- Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
- Partner with engineering leadership to develop a roadmap to improve reliability, scalability, operational excellence, and engineering efficiency across the Ads organization.
- Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
- Drive architecture reviews and influence technical decisions impacting critical revenue-generating systems.
- Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.
- Identify systemic reliability risks and drive long-term solutions that improve platform resilience.
- Establish reliability metrics around advertiser-critical user journeys such as campaign creation, ad delivery, auction participation, reporting, attribution, and billing.
- Mentor engineers and provide technical leadership across multiple teams.
View Full Description & ApplyYou'll be redirected to the employer's site