Senior DevOps Engineer, Infrastructure & Reliability

New
W
Worth AIComputer software
Workable workplace: remote; Workable locations: Orlando, Florida, United States. Atlanta, Georgia, United States. Tampa, Florida, United States. Miami, Florida, United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
8+ years in DevOps, SRE, or infrastructure engineering.
Required Skills
AWSPostgreSQLKafkaKubernetesCI/CDDevOpsTerraformDatadog

Requirements

  • Bring 8+ years of experience in DevOps, SRE, or infrastructure engineering.
  • Have proven experience designing and operating production Kubernetes environments at scale.
  • Bring deep hands-on expertise with AWS infrastructure and cloud networking.
  • Have strong experience building and maintaining Terraform modules across large cloud environments.
  • Demonstrate ownership of CI/CD systems and measurable improvement of DORA metrics.
  • Have experience leading incident response processes and driving postmortem outcomes.
  • Understand distributed systems, event-driven architectures such as Kafka, and database performance, including PostgreSQL.
  • Demonstrate the ability to modernize legacy infrastructure and eliminate manual operational toil.
  • Have a track record of taking scoped infrastructure projects from ambiguous beginnings to production without daily direction.
  • Build trust across teams while raising the reliability bar.

Responsibilities

  • Implement scalable Terraform infrastructure-as-code patterns to standardize cloud provisioning and reduce configuration drift.
  • Own and evolve the Kubernetes platform, ensuring workloads are secure, scalable, and resilient.
  • Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence.
  • Design and enforce secure networking, IAM, and secrets management strategies across environments.
  • Improve observability by refining metrics, logs, and tracing using DataDog.
  • Optimize cloud costs through rightsizing, autoscaling, and architectural improvements.
  • Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.
  • Refactor brittle or manually managed infrastructure into automated, testable, reproducible systems.
  • Introduce infrastructure tooling or architectural changes and support adoption through documentation, workshops, and hands-on assistance.
  • Partner with engineering teams to reduce friction in CI/CD, deployments, and cloud environments.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now