Site Reliability Engineer (Hadoop)
New
P
PulsePointHealth adtech
EU/UK/US employment, 9am–6pm ET US hoursFull-TimeSenior
Salary90,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- HadoopMicrosoft SQL ServerApache KafkaGrafanaPrometheusTerraformAnsible
Requirements
- 5+ years running distributed systems reliability at scale in production.
- Deep expertise in Kafka, Ceph, or similar distributed infrastructure.
- Proven ability to design for scale, reliability, and failure recovery.
- Experience mentoring engineers and making technical decisions.
- Willingness to work 9am–6pm ET US hours.
- Experience with Terraform, Ansible, and Puppet.
- Familiarity with monitoring tools like Prometheus, Grafana, and Icinga.
- Knowledge of ArgoCD and PagerDuty.
Responsibilities
- Ensure Kafka reliability through architecture, topic design, governance, and optimization.
- Manage Ceph reliability including operations, pool design, and capacity planning.
- Build operational automation to reduce manual toil and improve incident response.
- Support SQL Server backup and recovery pipelines.
- Build observability and self-service tooling for the data team.
View Full Description & ApplyYou'll be redirected to the employer's site