Senior DevOps Engineer
New
A
Akkadian LabsCollaboration Automation
RemoteFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of experience in DevOps or Site Reliability Engineering (SRE)
- Required Skills
- AWSDockerPythonBashKubernetesPrometheusCI/CDLinuxTerraformCloudFormation
Requirements
- 10+ years of experience in DevOps or Site Reliability Engineering (SRE).
- Expertise with AWS (e.g., EC2, ECS, S3, IAM, Lambda, CloudWatch).
- Expertise with infrastructure-as-code tools, including Terraform and CloudFormation.
- Strong knowledge of Linux environments (Rocky OS).
- Experienced with Docker and Kubernetes.
- Scripting ability in Python, Bash, or similar languages.
- Experience building or maintaining CI/CD pipelines.
- Experience in monitoring and observability tools such as Prometheus, Grafana, and ELK.
- Experience in implementing secure DevOps practices and compliance frameworks (SOC2, ISO).
- Experience supporting AI or machine learning workloads and compute environments.
- Experience supporting production systems and participating in on-call rotations.
Responsibilities
- Deploy and maintain scalable infrastructure in AWS and hybrid cloud environments.
- Manage infrastructure-as-code (IaC) using Terraform, CloudFormation, or similar tools.
- Design, deploy and manage AI agent workloads, including compute provisioning and resource scaling.
- Build and maintain model deployment pipelines for AI models.
- Monitor AI API consumption, infrastructure costs, and implement security guardrails.
- Manage monitoring and observability using Prometheus, Grafana, and the ELK stack.
- Build, maintain, and optimize CI/CD pipelines and automate operational tasks.
- Support compliance initiatives and vulnerability remediation efforts.
View Full Description & ApplyYou'll be redirected to the employer's site