Senior DevOps Engineer

New
B
BinagoraCloud infrastructure
Workable workplace: remote; Workable locations: ArgentinaContractSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of hands-on experience in DevOps/SRE roles managing production AWS environments
Required Skills
AWSDockerPythonJavaCI/CDTerraformCloudFormation

Requirements

  • Have 5+ years of hands-on experience in DevOps/SRE roles managing production AWS environments.
  • Have experience with AWS EC2, ECS/EKS, VPC, and RDS.
  • Demonstrate strong expertise with Docker containerization.
  • Use Terraform, CloudFormation, or AWS CDK for Infrastructure as Code.
  • Troubleshoot Python runtimes, including Gunicorn and Celery.
  • Troubleshoot Java runtimes, including JVM, thread pools, and garbage collection.
  • Have experience building CI/CD pipelines, automated testing, and safe deployment strategies.
  • Have hands-on proficiency with CloudWatch, OpenSearch, or Loki.
  • Understand disaster recovery concepts, including RTO/RPO.
  • Have Python and Bash scripting proficiency.
  • Nice to have: experience with advanced autoscaling based on queue depth, message age, or custom application metrics.
  • Nice to have: experience with capacity planning, sizing, and tuning CPU, memory, storage, and IOPS.
  • Nice to have: AWS Certified DevOps Engineer – Professional or Solutions Architect – Professional certification.

Responsibilities

  • Design and maintain multi-AZ AWS infrastructure, including networking, storage, and databases.
  • Build and operate Docker-based containerized applications and microservices in production.
  • Troubleshoot Python and Java applications, resolving memory, concurrency, and performance issues.
  • Manage Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
  • Develop CI/CD pipelines with integrated security scanning and testing.
  • Implement zero-downtime deployment strategies, including blue/green, canary, and rolling updates.
  • Centralize logging, metrics, tracing, and alerts using observability tools.
  • Configure application and cluster autoscaling based on load testing and usage metrics.
  • Implement resilience patterns such as circuit breakers, retries, and dead-letter queues.
  • Define disaster recovery plans and enforce security best practices across IAM and encryption.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now