Engineering Manager - Machine Learning

R
RecursionTechBio / Drug Discovery
This is a fully remote position based in Toronto, Canada.Full-TimeManager
Salary$210,070–$282,851 (CAD)
Apply NowOpens the employer's application page

Job Details

Required Skills
DockerPythonGCPKubernetesPyTorchMLOpsDistributed Systems

Requirements

  • Hands-on technical leadership or management experience in infrastructure, MLOps, and distributed systems.
  • Proven experience delivering impact in ML infrastructure, model deployment, distributed compute, and GPU optimization.
  • Ability to mentor, coach, and sponsor team members to drive technical and managerial growth.
  • Deep technical interest in machine learning, orchestration, and agentic systems.
  • Proficiency with current stack technologies such as Python, PyTorch, Docker, Kubernetes, and Ray.
  • Familiarity with MLOps and infrastructure tools like Weights & Biases, Prefect, BigQuery, Postgres, GCP, and CUDA.
  • Strong collaborative mindset for working across organizational boundaries.

Responsibilities

  • Lead and mentor a team of engineers focusing on ML infrastructure, distributed computing, and MLOps.
  • Build and operate platforms for training, deploying, and monitoring models across massive experimental datasets.
  • Partner with ML research and platform engineering teams to translate infrastructure requirements into scalable solutions.
  • Optimize GPU cluster utilization and implement orchestration for AI, ML, LLM, and agentic systems.
  • Establish company-wide MLOps standards to support rapid experimentation and reliable model deployment.
View Full Description & ApplyYou'll be redirected to the employer's site
$210,070–$282,851 (CAD)
Apply Now