Senior Platform Engineer
New
R
Rezilient HealthHealthcare technology
United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years in Site Reliability, DevOps, or infrastructure engineering roles
- Required Skills
- AWSDockerPythonGCPKubernetesAzurePrometheusCI/CDTerraformDatadog
Requirements
- Bachelor's degree in computer science, software engineering, or a related field, or equivalent hands-on experience.
- 5+ years in Site Reliability, DevOps, or infrastructure engineering roles at startup or growth-stage organizations.
- Hands-on experience operating production systems on AWS, GCP, or Azure, including compute, networking, storage, and managed database services.
- Proficiency with Infrastructure as Code, such as Terraform, and configuration management.
- Production experience with Docker and Kubernetes, including deployment, scaling, and troubleshooting.
- Experience building CI/CD pipelines and release automation.
- Hands-on experience with observability and monitoring tools such as Datadog, Prometheus, Grafana, CloudWatch, or ELK/OpenSearch, and defining SLOs, SLIs, and error budgets.
- Scripting and automation skills in Python, Go, or Bash, and fluency with command-line and cloud CLIs.
- Experience leading incident response and on-call, including postmortem and root-cause analysis practices.
- Understanding of infrastructure and network security, secrets management, and encryption; familiarity with sensitive or protected data, with PHI/HIPAA experience strongly preferred.
- Proficiency with Git and version control workflows; familiarity with the Agile Development Framework and ideally Jira and Confluence.
Responsibilities
- Design, provision, and maintain cloud infrastructure across development, staging, and production using Infrastructure as Code.
- Own CI/CD pipelines, automated testing gates, blue/green and canary rollouts, and rollbacks.
- Build and operate metrics, logging, distributed tracing, dashboards, and alerting; define SLOs, SLIs, and error budgets.
- Lead on-call incident response, including triage, mitigation, communication, and blameless postmortems.
- Automate operational toil through scripting and tooling.
- Design capacity planning, load and failure testing, autoscaling, redundancy, disaster recovery, and backup strategies.
- Manage Docker and Kubernetes workloads, including networking, service discovery, and resource management.
- Partner with Security and Engineering to harden infrastructure through secrets management, network segmentation, encryption, vulnerability scanning, patching, and audit logging.
- Collaborate with development teams to embed reliability and operability into services from design through production.
View Full Description & ApplyYou'll be redirected to the employer's site