Machine Learning DevOps - Cloud and Compute Cluster
New
P
PathwayArtificial Intelligence
Candidates based anywhere in the EU, United States, and Canada will be considered.Full-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSDockerPythonGCPKubernetesAzureCI/CDLinuxTerraform
Requirements
- BSc in Computer Science or Information Technology.
- Familiarity with Linux, shell scripts, and cluster configuration scripts.
- Proficiency in workload management, containerization and orchestration (Slurm, Docker, Kubernetes).
- Solid grasp of CI/CD tools and workflows (GitHub Actions, Jenkins, Gitlab CI, etc.).
- Cloud infrastructure knowledge (AWS, GCP, Azure), especially in ML services.
- Familiarity with monitoring/logging tools (Grafana, CloudWatch, Prometheus, Loki).
- Experience with infrastructure as code (Terraform, CloudFormation, cluster-toolkit).
- Experience with ML pipeline orchestration tools (e.g., MLflow, Kubeflow, Airflow, Metaflow).
- Programming skills in Python, including ML libraries like TensorFlow and PyTorch.
- Experience with cluster, systems, and networks administration.
Responsibilities
- Optimize infrastructure for ML training and inference (e.g., GPUs, distributed compute).
- Automate and maintain ML/LLM pipelines (data ingestion, training, validation, deployment).
- Manage model versioning, reproducibility, and traceability.
- Work with terabyte-large datasets.
- Implement ML-centric CI/CD practices.
- Monitor model performance and data drift in production.
- Collaborate with machine learning engineers, software engineers, and platform teams.
View Full Description & ApplyYou'll be redirected to the employer's site