Senior DevOps Engineer

New
A
Akkadian LabsCollaboration Automation
RemoteFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
10+ years of experience in DevOps or Site Reliability Engineering (SRE)
Required Skills
AWSDockerPythonBashKubernetesPrometheusCI/CDLinuxTerraformCloudFormation

Requirements

  • 10+ years of experience in DevOps or Site Reliability Engineering (SRE).
  • Expertise with AWS (e.g., EC2, ECS, S3, IAM, Lambda, CloudWatch).
  • Expertise with infrastructure-as-code tools, including Terraform and CloudFormation.
  • Strong knowledge of Linux environments (Rocky OS).
  • Experienced with Docker and Kubernetes.
  • Scripting ability in Python, Bash, or similar languages.
  • Experience building or maintaining CI/CD pipelines.
  • Experience in monitoring and observability tools such as Prometheus, Grafana, and ELK.
  • Experience in implementing secure DevOps practices and compliance frameworks (SOC2, ISO).
  • Experience supporting AI or machine learning workloads and compute environments.
  • Experience supporting production systems and participating in on-call rotations.

Responsibilities

  • Deploy and maintain scalable infrastructure in AWS and hybrid cloud environments.
  • Manage infrastructure-as-code (IaC) using Terraform, CloudFormation, or similar tools.
  • Design, deploy and manage AI agent workloads, including compute provisioning and resource scaling.
  • Build and maintain model deployment pipelines for AI models.
  • Monitor AI API consumption, infrastructure costs, and implement security guardrails.
  • Manage monitoring and observability using Prometheus, Grafana, and the ELK stack.
  • Build, maintain, and optimize CI/CD pipelines and automate operational tasks.
  • Support compliance initiatives and vulnerability remediation efforts.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now