Senior DevOps Engineer, Infrastructure & Reliability
New
W
Worth AIComputer software
Workable workplace: remote; Workable locations: Orlando, Florida, United States. Atlanta, Georgia, United States. Tampa, Florida, United States. Miami, Florida, United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years in DevOps, SRE, or infrastructure engineering.
- Required Skills
- AWSPostgreSQLKafkaKubernetesCI/CDDevOpsTerraformDatadog
Requirements
- Bring 8+ years of experience in DevOps, SRE, or infrastructure engineering.
- Have proven experience designing and operating production Kubernetes environments at scale.
- Bring deep hands-on expertise with AWS infrastructure and cloud networking.
- Have strong experience building and maintaining Terraform modules across large cloud environments.
- Demonstrate ownership of CI/CD systems and measurable improvement of DORA metrics.
- Have experience leading incident response processes and driving postmortem outcomes.
- Understand distributed systems, event-driven architectures such as Kafka, and database performance, including PostgreSQL.
- Demonstrate the ability to modernize legacy infrastructure and eliminate manual operational toil.
- Have a track record of taking scoped infrastructure projects from ambiguous beginnings to production without daily direction.
- Build trust across teams while raising the reliability bar.
Responsibilities
- Implement scalable Terraform infrastructure-as-code patterns to standardize cloud provisioning and reduce configuration drift.
- Own and evolve the Kubernetes platform, ensuring workloads are secure, scalable, and resilient.
- Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence.
- Design and enforce secure networking, IAM, and secrets management strategies across environments.
- Improve observability by refining metrics, logs, and tracing using DataDog.
- Optimize cloud costs through rightsizing, autoscaling, and architectural improvements.
- Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.
- Refactor brittle or manually managed infrastructure into automated, testable, reproducible systems.
- Introduce infrastructure tooling or architectural changes and support adoption through documentation, workshops, and hands-on assistance.
- Partner with engineering teams to reduce friction in CI/CD, deployments, and cloud environments.
View Full Description & ApplyYou'll be redirected to the employer's site