Senior Site Reliability Engineer
New
B
BloomreachAI Personalization
Working from one of our Central European offices (Bratislava, Prague, or Brno), or remotely (Czechia, Slovakia)Full-TimeSenior
Salary€41.600 — €52.000 EUR
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PostgreSQLPythonElasticSearchGCPKafkaKubernetesGoGrafanaPrometheus
Requirements
- Proven experience operating services on Kubernetes in a major cloud environment, preferably GCP.
- Strong experience with observability and incident diagnosis for distributed systems.
- Proficiency in Go or Python.
- Practical experience operating relational databases, preferably PostgreSQL or Cloud SQL.
- Experience with large-scale data systems such as Bigtable, Elasticsearch, or Kafka.
- Experience designing or operating asynchronous job-processing systems and queues.
- Strong knowledge of CI/CD, Infrastructure as Code, and deployment automation.
- Ability to communicate technical risks and trade-offs to diverse stakeholders.
- Comfort participating in on-call rotations and responding to production incidents.
- Experience with API reliability, rate limiting, and multi-tenant isolation.
Responsibilities
- Own and improve the reliability posture of Fuse services, workers, APIs, queues, storage, and synchronization pipelines.
- Establish meaningful SLIs, SLOs, and error budgets for critical customer-facing systems.
- Build end-to-end observability across Data Hub item collections using tools like Grafana and Prometheus.
- Improve deployment automation, rollbacks, and release safety for Kubernetes-based infrastructure.
- Lead incident investigations, L3 support rotations, and root-cause analysis for distributed failures.
- Partner with engineering teams to design scalable, observable, and cost-aware system architectures.
- Enforce security, isolation, and compliance controls including ISO and SOC 2 requirements.
View Full Description & ApplyYou'll be redirected to the employer's site