Senior DevOps Engineer, Infrastructure & Reliability

New
W
Worth AIComputer software
Orlando, Florida, United States. Atlanta, Georgia, United States. Tampa, Florida, United States. Miami, Florida, United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
8+ years in DevOps, SRE, or infrastructure engineering
Required Skills
AWSPostgreSQLPythonBashKafkaKubernetesCI/CDTerraformDatadog

Requirements

  • 8+ years of experience in DevOps, SRE, or infrastructure engineering.
  • Proven experience designing and operating production Kubernetes environments at scale.
  • Deep hands-on expertise with AWS infrastructure and cloud networking.
  • Strong experience building and maintaining Terraform modules across large cloud environments.
  • Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics.
  • Experience leading incident response processes and driving meaningful postmortem outcomes.
  • Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
  • Proven ability to modernize legacy infrastructure and eliminate manual operational toil.
  • Track record of taking a scoped infrastructure project from an ambiguous starting point to production.
  • Demonstrated ability to build trust across teams while raising the reliability bar.
  • Willingness to travel to Orlando, Florida at least twice per year for Town Halls, team collaboration, and orientation.

Responsibilities

  • Implement scalable Infrastructure-as-Code patterns using Terraform to standardize cloud provisioning.
  • Own and evolve the Kubernetes platform, ensuring workloads are secure, scalable, and resilient.
  • Optimize CI/CD pipelines to improve deployment frequency and lead times.
  • Design and enforce secure networking, IAM, and secrets management strategies.
  • Refine observability metrics, logs, and tracing using DataDog to ensure actionable insights.
  • Drive cloud cost efficiency through rightsizing and autoscaling strategies.
  • Implement disaster recovery, backup strategies, and multi-region resilience initiatives.
  • Refactor legacy infrastructure into automated, testable, and reproducible systems.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now