Staff Site Reliability Engineer, Ads
New
J
JobgetherAdvertising Technology
Flexible-first remote work environment within the United States.Full-TimeStaff
SalaryBase salary range of $217,000–$303,900 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- KafkaKubernetesClickhouseGoSparkBigQueryDistributed Systems
Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related discipline, with experience operating large-scale distributed systems.
- Proven experience evolving and supporting high-traffic, user-facing production environments with demanding availability and performance requirements.
- Deep expertise in distributed systems, scalability engineering, cloud-native architectures, and highly available system design.
- Strong software engineering capabilities, ideally with experience in backend programming languages such as Go.
- Extensive knowledge of observability practices and technologies, including metrics, logging, tracing, alerting, and performance monitoring.
- Demonstrated experience improving reliability through SLOs, automation, incident management, performance optimization, and systematic operational practices.
- Strong troubleshooting and problem-solving abilities across complex, modern distributed technology stacks.
- Excellent cross-functional communication and collaboration skills, with the ability to influence technical direction and align teams around shared reliability goals.
- Experience supporting advertising technology or other large-scale, revenue-critical platforms is highly desirable.
- Hands-on experience with Kubernetes, cloud infrastructure, and large-scale distributed systems is beneficial.
Responsibilities
- Lead reliability initiatives across critical advertising domains, including ad serving, auctions, targeting, reporting, measurement, attribution, and billing.
- Partner with engineering leadership to establish and execute roadmaps focused on reliability, scalability, operational excellence, and developer productivity.
- Design and build scalable platforms, tooling, automation, and infrastructure capabilities that improve system resilience and engineering efficiency.
- Lead architecture reviews and influence technical decisions for high-traffic, revenue-critical distributed systems.
- Establish and monitor reliability metrics and SLOs around critical advertiser and platform journeys, using data to identify risks and prioritize improvements.
- Participate in on-call rotations, lead complex incident investigations, and coordinate cross-functional responses to major production events.
- Identify systemic reliability risks and implement durable solutions that improve availability, performance, scalability, and operational maturity.
- Drive automation, observability, incident management, performance optimization, and other practices that strengthen production reliability.
- Mentor engineers and provide technical leadership across multiple teams, helping raise engineering standards and reliability expertise.
- Collaborate with Product, Data Science, Infrastructure, and Engineering stakeholders to ensure reliability considerations are embedded into product and infrastructure investments.
View Full Description & ApplyYou'll be redirected to the employer's site