Machine Learning DevOps - Cloud and Compute Cluster

New
P
PathwayArtificial Intelligence
Candidates based anywhere in the EU, United States, and Canada will be considered.Full-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSDockerPythonGCPKubernetesAzureCI/CDLinuxTerraform

Requirements

  • BSc in Computer Science or Information Technology.
  • Familiarity with Linux, shell scripts, and cluster configuration scripts.
  • Proficiency in workload management, containerization and orchestration (Slurm, Docker, Kubernetes).
  • Solid grasp of CI/CD tools and workflows (GitHub Actions, Jenkins, Gitlab CI, etc.).
  • Cloud infrastructure knowledge (AWS, GCP, Azure), especially in ML services.
  • Familiarity with monitoring/logging tools (Grafana, CloudWatch, Prometheus, Loki).
  • Experience with infrastructure as code (Terraform, CloudFormation, cluster-toolkit).
  • Experience with ML pipeline orchestration tools (e.g., MLflow, Kubeflow, Airflow, Metaflow).
  • Programming skills in Python, including ML libraries like TensorFlow and PyTorch.
  • Experience with cluster, systems, and networks administration.

Responsibilities

  • Optimize infrastructure for ML training and inference (e.g., GPUs, distributed compute).
  • Automate and maintain ML/LLM pipelines (data ingestion, training, validation, deployment).
  • Manage model versioning, reproducibility, and traceability.
  • Work with terabyte-large datasets.
  • Implement ML-centric CI/CD practices.
  • Monitor model performance and data drift in production.
  • Collaborate with machine learning engineers, software engineers, and platform teams.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now