Staff Site Reliability Engineer, Ads

New
J
JobgetherAdvertising Technology
Flexible-first remote work environment within the United States.Full-TimeStaff
SalaryBase salary range of $217,000–$303,900 USD
Apply NowOpens the employer's application page

Job Details

Experience
8+ years
Required Skills
KafkaKubernetesClickhouseGoSparkBigQueryDistributed Systems

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related discipline, with experience operating large-scale distributed systems.
  • Proven experience evolving and supporting high-traffic, user-facing production environments with demanding availability and performance requirements.
  • Deep expertise in distributed systems, scalability engineering, cloud-native architectures, and highly available system design.
  • Strong software engineering capabilities, ideally with experience in backend programming languages such as Go.
  • Extensive knowledge of observability practices and technologies, including metrics, logging, tracing, alerting, and performance monitoring.
  • Demonstrated experience improving reliability through SLOs, automation, incident management, performance optimization, and systematic operational practices.
  • Strong troubleshooting and problem-solving abilities across complex, modern distributed technology stacks.
  • Excellent cross-functional communication and collaboration skills, with the ability to influence technical direction and align teams around shared reliability goals.
  • Experience supporting advertising technology or other large-scale, revenue-critical platforms is highly desirable.
  • Hands-on experience with Kubernetes, cloud infrastructure, and large-scale distributed systems is beneficial.

Responsibilities

  • Lead reliability initiatives across critical advertising domains, including ad serving, auctions, targeting, reporting, measurement, attribution, and billing.
  • Partner with engineering leadership to establish and execute roadmaps focused on reliability, scalability, operational excellence, and developer productivity.
  • Design and build scalable platforms, tooling, automation, and infrastructure capabilities that improve system resilience and engineering efficiency.
  • Lead architecture reviews and influence technical decisions for high-traffic, revenue-critical distributed systems.
  • Establish and monitor reliability metrics and SLOs around critical advertiser and platform journeys, using data to identify risks and prioritize improvements.
  • Participate in on-call rotations, lead complex incident investigations, and coordinate cross-functional responses to major production events.
  • Identify systemic reliability risks and implement durable solutions that improve availability, performance, scalability, and operational maturity.
  • Drive automation, observability, incident management, performance optimization, and other practices that strengthen production reliability.
  • Mentor engineers and provide technical leadership across multiple teams, helping raise engineering standards and reliability expertise.
  • Collaborate with Product, Data Science, Infrastructure, and Engineering stakeholders to ensure reliability considerations are embedded into product and infrastructure investments.
View Full Description & ApplyYou'll be redirected to the employer's site
Base salary range of $217,000–$303,900 USD
Apply Now