Senior DevOps Engineer, Reliability & Platform Operations
M
MedTrainerHealthcare technology
100% remote work from anywhere in Mexico.Full-TimeSenior
Salary70,000 - 90,000 MXN per month
Apply NowOpens the employer's application page
Job Details
- Languages
- Strong written and verbal English communication skills.
- Experience
- 8+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, cloud operations, infrastructure automation, production operations, or related roles
- Required Skills
- PythonKubernetesMySQLAzureLinuxTerraformAnsibleGitHub Actions
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent professional experience.
- 8+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, cloud operations, infrastructure automation, production operations, or related roles.
- Strong hands-on experience with Azure and production Kubernetes / AKS environments.
- Experience designing or improving CI/CD, infrastructure as code, configuration management, observability, and deployment automation at scale.
- Strong production troubleshooting skills across cloud infrastructure, Linux, containers, Kubernetes, databases, middleware, CI/CD pipelines, and application-supporting services.
- Ability to understand application behavior and support production reliability of PHP Symfony applications.
- Experience operating or supporting MySQL, ProxySQL, and RabbitMQ in production environments.
- Experience with Terraform or Pulumi; Pulumi with Python is preferred.
- Experience with Ansible; Python is required and Bash is preferred.
- Strong understanding of secure delivery and infrastructure security, including secrets management, least privilege, access reviews, CI/CD security, and audit support.
- Experience with backup, restore, disaster recovery, business continuity, capacity planning, performance tuning, and cost optimization.
- Ability to mentor engineers, define technical standards, and lead technical initiatives without formal people-management authority.
- Strong written and verbal English communication skills.
Responsibilities
- Improve reliability practices across production systems, including SLIs, SLOs, error budgets, postmortems, runbooks, and corrective-action tracking.
- Lead reliability improvements across Azure cloud platforms, AKS, CI/CD pipelines, application environments, databases, and middleware.
- Operate and improve AKS clusters, including upgrades, autoscaling, node pools, networking, storage, identity, and security.
- Improve GitHub Actions workflows, reusable automation, secure deployment patterns, rollback support, and pipeline optimization.
- Build automation, self-service capabilities, reusable infrastructure patterns, and operational procedures to reduce toil and infrastructure drift.
- Implement and maintain infrastructure as code and configuration management, including Ansible-based automation.
- Improve observability across metrics, logs, traces, dashboards, alerting, APM, capacity planning, and incident dashboards.
- Support production incident response through triage coordination, mitigation, postmortems, and corrective actions.
- Operate, tune, monitor, back up, troubleshoot, and support MySQL, ProxySQL, and RabbitMQ.
- Strengthen infrastructure and delivery security, backup and disaster recovery, cloud cost visibility, and operational resilience.
View Full Description & ApplyYou'll be redirected to the employer's site