Lead Site Reliability Engineer

C
CloudLinuxLinux server security
Warsaw, Masovian Voivodeship, Poland. Bucharest, Bucharest, Romania. Sofia, Sofia City Province, Bulgaria. Belgrade, Vojvodina, Serbia. Yerevan, Yerevan, Armenia. Tbilisi, Tbilisi, Georgia, UTC−5 … UTC+8Full-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonClickhouseGoGrafanaPrometheusRustAnsibleDistributed Systems

Requirements

  • Substantial production-engineering or SRE experience, including defining SLO frameworks.
  • Strong proficiency in Python.
  • Ability to read and modify Go or Rust code.
  • Deep practical grip on telemetry stacks like Prometheus, Grafana, and OpenMetrics.
  • Experience with columnar stores for high-cardinality data, such as ClickHouse.
  • Expertise in distributed systems debugging on bare metal and long-lived hosts.
  • Proficiency in configuration management and CI tools like Ansible and GitLab CI or Jenkins.
  • Strong written communication skills for effective async collaboration.
  • Experience with designing measurement systems for remote or edge environments.

Responsibilities

  • Define service level indicators (SLIs) for ~70 product components in collaboration with squad leads.
  • Develop a robust telemetry collection pipeline that is push-based, sampled, and privacy-constrained.
  • Implement symptom-based, SLO-anchored alerting systems to minimize non-actionable alerts.
  • Establish clear escalation paths and ownership maps for engineering squads.
  • Coach teams on incident command practices, blameless postmortems, and on-call management.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now