Senior DevOps Engineer, Infrastructure & Reliability
New
W
Worth AIComputer software
Orlando, Florida, United States. Atlanta, Georgia, United States. Tampa, Florida, United States. Miami, Florida, United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years in DevOps, SRE, or infrastructure engineering
- Required Skills
- AWSPostgreSQLPythonBashKafkaKubernetesCI/CDTerraformDatadog
Requirements
- 8+ years of experience in DevOps, SRE, or infrastructure engineering.
- Proven experience designing and operating production Kubernetes environments at scale.
- Deep hands-on expertise with AWS infrastructure and cloud networking.
- Strong experience building and maintaining Terraform modules across large cloud environments.
- Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics.
- Experience leading incident response processes and driving meaningful postmortem outcomes.
- Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
- Proven ability to modernize legacy infrastructure and eliminate manual operational toil.
- Track record of taking a scoped infrastructure project from an ambiguous starting point to production.
- Demonstrated ability to build trust across teams while raising the reliability bar.
- Willingness to travel to Orlando, Florida at least twice per year for Town Halls, team collaboration, and orientation.
Responsibilities
- Implement scalable Infrastructure-as-Code patterns using Terraform to standardize cloud provisioning.
- Own and evolve the Kubernetes platform, ensuring workloads are secure, scalable, and resilient.
- Optimize CI/CD pipelines to improve deployment frequency and lead times.
- Design and enforce secure networking, IAM, and secrets management strategies.
- Refine observability metrics, logs, and tracing using DataDog to ensure actionable insights.
- Drive cloud cost efficiency through rightsizing and autoscaling strategies.
- Implement disaster recovery, backup strategies, and multi-region resilience initiatives.
- Refactor legacy infrastructure into automated, testable, and reproducible systems.
View Full Description & ApplyYou'll be redirected to the employer's site