Lead Site Reliability Engineer
C
CloudLinuxLinux server security
Warsaw, Masovian Voivodeship, Poland. Bucharest, Bucharest, Romania. Sofia, Sofia City Province, Bulgaria. Belgrade, Vojvodina, Serbia. Yerevan, Yerevan, Armenia. Tbilisi, Tbilisi, Georgia, UTC−5 … UTC+8Full-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonClickhouseGoGrafanaPrometheusRustAnsibleDistributed Systems
Requirements
- Substantial production-engineering or SRE experience, including defining SLO frameworks.
- Strong proficiency in Python.
- Ability to read and modify Go or Rust code.
- Deep practical grip on telemetry stacks like Prometheus, Grafana, and OpenMetrics.
- Experience with columnar stores for high-cardinality data, such as ClickHouse.
- Expertise in distributed systems debugging on bare metal and long-lived hosts.
- Proficiency in configuration management and CI tools like Ansible and GitLab CI or Jenkins.
- Strong written communication skills for effective async collaboration.
- Experience with designing measurement systems for remote or edge environments.
Responsibilities
- Define service level indicators (SLIs) for ~70 product components in collaboration with squad leads.
- Develop a robust telemetry collection pipeline that is push-based, sampled, and privacy-constrained.
- Implement symptom-based, SLO-anchored alerting systems to minimize non-actionable alerts.
- Establish clear escalation paths and ownership maps for engineering squads.
- Coach teams on incident command practices, blameless postmortems, and on-call management.
View Full Description & ApplyYou'll be redirected to the employer's site