Cloud Engineer (Azure Platform Engineer)
New
S
Smile Digital HealthHealth Data
Remote OntarioFull-TimeMiddle
Salary115,000 - 130,000 CAD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of hands-on Kubernetes experience, including at least 2 years running AKS in production.
- Required Skills
- PythonBashKubernetesGrafanaPrometheusTerraformAzure DevOpsHelm
Requirements
- 5+ years of hands-on Kubernetes experience, including at least 2 years running AKS in production.
- Strong knowledge of Kubernetes internals including scheduling, networking, storage, and RBAC.
- Proficiency in Azure CNI networking and AKS private cluster configuration.
- Hands-on experience integrating AKS with Azure PaaS services and troubleshooting network dependencies.
- Experience with Helm and GitOps workflows (Flux/ArgoCD).
- Working knowledge of Terraform for infrastructure as code.
- Working knowledge of Azure DevOps or similar CI/CD tooling.
- Hands-on experience with Grafana (Prometheus, Loki, Tempo) and Azure-native monitoring.
- Experience with container/cluster vulnerability management.
- Scripting skills in Bash or Python.
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) certification required.
- Microsoft Certified: Azure Administrator (AZ-104) required.
Responsibilities
- Act as the subject-matter expert (SME) for Kubernetes deployments, troubleshooting, and production issues across all environments.
- Design and deploy Azure Kubernetes Service (AKS) clusters with private cluster configurations and RBAC.
- Manage the health, scaling, and lifecycle of production AKS clusters, including upgrades and node pool management.
- Configure AKS networking and integrate with Azure PaaS services like ACR, Key Vault, and Azure SQL.
- Manage containerized deployments using Docker and Helm while enforcing security policies.
- Implement infrastructure as code using Terraform and manage CI/CD pipelines with GitOps workflows.
- Design and maintain observability across the Grafana stack and Azure-native monitoring tools.
- Participate in on-call rotations and lead root-cause analysis for production incidents.
View Full Description & ApplyYou'll be redirected to the employer's site