Engineering Manager - Machine Learning
R
RecursionTechBio / Drug Discovery
This is a fully remote position based in Toronto, Canada.Full-TimeManager
Salary$210,070–$282,851 (CAD)
Apply NowOpens the employer's application page
Job Details
- Required Skills
- DockerPythonGCPKubernetesPyTorchMLOpsDistributed Systems
Requirements
- Hands-on technical leadership or management experience in infrastructure, MLOps, and distributed systems.
- Proven experience delivering impact in ML infrastructure, model deployment, distributed compute, and GPU optimization.
- Ability to mentor, coach, and sponsor team members to drive technical and managerial growth.
- Deep technical interest in machine learning, orchestration, and agentic systems.
- Proficiency with current stack technologies such as Python, PyTorch, Docker, Kubernetes, and Ray.
- Familiarity with MLOps and infrastructure tools like Weights & Biases, Prefect, BigQuery, Postgres, GCP, and CUDA.
- Strong collaborative mindset for working across organizational boundaries.
Responsibilities
- Lead and mentor a team of engineers focusing on ML infrastructure, distributed computing, and MLOps.
- Build and operate platforms for training, deploying, and monitoring models across massive experimental datasets.
- Partner with ML research and platform engineering teams to translate infrastructure requirements into scalable solutions.
- Optimize GPU cluster utilization and implement orchestration for AI, ML, LLM, and agentic systems.
- Establish company-wide MLOps standards to support rapid experimentation and reliable model deployment.
View Full Description & ApplyYou'll be redirected to the employer's site