Senior DevOps Engineer, AI Platform
New
J
JobgetherAI Platform/DevOps
Full-time, fully remote position available across Canada.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- PostgreSQLPythonKubernetesRabbitmqAzureRedisCI/CDLinuxTerraform
Requirements
- 7+ years of professional experience in DevOps, SRE, or Platform Engineering.
- Strong hands-on experience operating production Kubernetes environments (networking, scheduling, security).
- Strong Microsoft Azure experience (AKS, identity, networking, monitoring); Oracle Cloud experience preferred.
- Deep understanding of cloud networking (subnets, routing, NAT, firewalls, TLS, DNS).
- Proven experience with Jenkins, Bitbucket, Docker, Terraform, Helm, and Kubernetes.
- Experience supporting web applications and backend services (REST APIs, microservices, asynchronous architectures).
- Knowledge of databases, caching, and messaging (PostgreSQL, Redis, RabbitMQ).
- Experience with observability tools (OpenTelemetry, Grafana, Prometheus, Sentry).
- Strong Linux systems administration and production troubleshooting skills.
- Working knowledge of Python (FastAPI); familiarity with C#, Java, Go, JS, or TS.
- Experience supporting AI/ML platforms, LLM gateways, or RAG pipelines is highly desirable.
Responsibilities
- Design, provision, and troubleshoot Kubernetes environments on Azure Kubernetes Service (AKS) and Oracle Kubernetes Engine.
- Build and support infrastructure for AI workloads including LLM gateways, Python-based agent runtimes, and RAG pipelines.
- Manage cloud networking components including virtual networks, load balancers, DNS, TLS, and private connectivity.
- Build and maintain CI/CD pipelines using Jenkins, Bitbucket, Docker, and ArgoCD.
- Automate infrastructure provisioning using Terraform, Helm, and Infrastructure as Code practices.
- Implement comprehensive observability and monitoring using metrics, logs, and distributed tracing.
- Own production reliability, incident response, root cause analysis, and infrastructure cost optimization.
View Full Description & ApplyYou'll be redirected to the employer's site