Senior DevOps Engineer, AI Platform

New
J
JobgetherAI Platform/DevOps
Full-time, fully remote position available across Canada.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
7+ years
Required Skills
PostgreSQLPythonKubernetesRabbitmqAzureRedisCI/CDLinuxTerraform

Requirements

  • 7+ years of professional experience in DevOps, SRE, or Platform Engineering.
  • Strong hands-on experience operating production Kubernetes environments (networking, scheduling, security).
  • Strong Microsoft Azure experience (AKS, identity, networking, monitoring); Oracle Cloud experience preferred.
  • Deep understanding of cloud networking (subnets, routing, NAT, firewalls, TLS, DNS).
  • Proven experience with Jenkins, Bitbucket, Docker, Terraform, Helm, and Kubernetes.
  • Experience supporting web applications and backend services (REST APIs, microservices, asynchronous architectures).
  • Knowledge of databases, caching, and messaging (PostgreSQL, Redis, RabbitMQ).
  • Experience with observability tools (OpenTelemetry, Grafana, Prometheus, Sentry).
  • Strong Linux systems administration and production troubleshooting skills.
  • Working knowledge of Python (FastAPI); familiarity with C#, Java, Go, JS, or TS.
  • Experience supporting AI/ML platforms, LLM gateways, or RAG pipelines is highly desirable.

Responsibilities

  • Design, provision, and troubleshoot Kubernetes environments on Azure Kubernetes Service (AKS) and Oracle Kubernetes Engine.
  • Build and support infrastructure for AI workloads including LLM gateways, Python-based agent runtimes, and RAG pipelines.
  • Manage cloud networking components including virtual networks, load balancers, DNS, TLS, and private connectivity.
  • Build and maintain CI/CD pipelines using Jenkins, Bitbucket, Docker, and ArgoCD.
  • Automate infrastructure provisioning using Terraform, Helm, and Infrastructure as Code practices.
  • Implement comprehensive observability and monitoring using metrics, logs, and distributed tracing.
  • Own production reliability, incident response, root cause analysis, and infrastructure cost optimization.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now