Site Reliability Engineer (Hadoop)

New
P
PulsePointHealth adtech
EU/UK/US employment, 9am–6pm ET US hoursFull-TimeSenior
Salary90,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
HadoopMicrosoft SQL ServerApache KafkaGrafanaPrometheusTerraformAnsible

Requirements

  • 5+ years running distributed systems reliability at scale in production.
  • Deep expertise in Kafka, Ceph, or similar distributed infrastructure.
  • Proven ability to design for scale, reliability, and failure recovery.
  • Experience mentoring engineers and making technical decisions.
  • Willingness to work 9am–6pm ET US hours.
  • Experience with Terraform, Ansible, and Puppet.
  • Familiarity with monitoring tools like Prometheus, Grafana, and Icinga.
  • Knowledge of ArgoCD and PagerDuty.

Responsibilities

  • Ensure Kafka reliability through architecture, topic design, governance, and optimization.
  • Manage Ceph reliability including operations, pool design, and capacity planning.
  • Build operational automation to reduce manual toil and improve incident response.
  • Support SQL Server backup and recovery pipelines.
  • Build observability and self-service tooling for the data team.
View Full Description & ApplyYou'll be redirected to the employer's site
90,000 - 150,000 USD per year
Apply Now