Site Reliability Engineer
New
P
PulsePointHealth and AdTech
Worldwide, 9am–6pm ETFull-TimeMiddle
Salary90,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Required Skills
- KafkaKubernetesPrometheusRedisTerraform
Requirements
- Experience operating production infrastructure at meaningful scale.
- Understanding of how distributed systems fail and recover.
- Proficiency with Kubernetes internals, expansion, and tuning.
- Experience with bare-metal infrastructure.
- Preference for automation over manual operational work.
- Ability to simplify systems to reduce complexity.
- Ownership extending beyond the boundaries of a single component.
- Willingness to work 9am–6pm ET US hours.
Responsibilities
- Design, build, and operate the Kubernetes platform architecture and lifecycle management.
- Own reliability, observability, and incident response across platform services.
- Build infrastructure automation and GitOps workflows to reduce operational toil.
- Manage networking, service connectivity, and platform security.
- Improve developer experience through self-service platform capabilities.
- Operate large-scale distributed systems running on bare-metal infrastructure.
View Full Description & ApplyYou'll be redirected to the employer's site