Senior DevOps Engineer
New
B
BinagoraCloud infrastructure
Workable workplace: remote; Workable locations: ArgentinaContractSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of hands-on experience in DevOps/SRE roles managing production AWS environments
- Required Skills
- AWSDockerPythonJavaCI/CDTerraformCloudFormation
Requirements
- Have 5+ years of hands-on experience in DevOps/SRE roles managing production AWS environments.
- Have experience with AWS EC2, ECS/EKS, VPC, and RDS.
- Demonstrate strong expertise with Docker containerization.
- Use Terraform, CloudFormation, or AWS CDK for Infrastructure as Code.
- Troubleshoot Python runtimes, including Gunicorn and Celery.
- Troubleshoot Java runtimes, including JVM, thread pools, and garbage collection.
- Have experience building CI/CD pipelines, automated testing, and safe deployment strategies.
- Have hands-on proficiency with CloudWatch, OpenSearch, or Loki.
- Understand disaster recovery concepts, including RTO/RPO.
- Have Python and Bash scripting proficiency.
- Nice to have: experience with advanced autoscaling based on queue depth, message age, or custom application metrics.
- Nice to have: experience with capacity planning, sizing, and tuning CPU, memory, storage, and IOPS.
- Nice to have: AWS Certified DevOps Engineer – Professional or Solutions Architect – Professional certification.
Responsibilities
- Design and maintain multi-AZ AWS infrastructure, including networking, storage, and databases.
- Build and operate Docker-based containerized applications and microservices in production.
- Troubleshoot Python and Java applications, resolving memory, concurrency, and performance issues.
- Manage Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
- Develop CI/CD pipelines with integrated security scanning and testing.
- Implement zero-downtime deployment strategies, including blue/green, canary, and rolling updates.
- Centralize logging, metrics, tracing, and alerts using observability tools.
- Configure application and cluster autoscaling based on load testing and usage metrics.
- Implement resilience patterns such as circuit breakers, retries, and dead-letter queues.
- Define disaster recovery plans and enforce security best practices across IAM and encryption.
View Full Description & ApplyYou'll be redirected to the employer's site