MLOps Engineer

New
J
JobgetherMachine Learning
Germany; Fully remote work within EuropeFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
At least 5 years
Required Skills
PythonKubeflowKubernetesMLFlowPyTorchGoTensorflowCI/CDTerraformMLOps

Requirements

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
  • At least 5 years of experience in MLOps, DevOps, or related engineering roles supporting production machine learning environments.
  • Proven experience designing and building MLOps infrastructure from the ground up using platforms such as MLflow, Weights & Biases, Kubeflow, or similar.
  • Strong hands-on experience with machine learning frameworks including PyTorch and TensorFlow.
  • Experience with model serving technologies such as TorchServe, TensorFlow Serving, Triton, or KServe.
  • Solid experience developing and managing scalable data pipelines and Kubernetes environments.
  • Experience with cloud infrastructure (AWS, GCP, or Azure) and Infrastructure as Code solutions including Terraform, Helm, or GitOps.
  • Strong programming skills in Python, Bash, and Go, with a focus on maintainable, scalable, and production-quality software.
  • Knowledge of AI system security, model governance, compliance, monitoring, and observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.

Responsibilities

  • Develop, automate, and maintain scalable machine learning pipelines, CI/CD workflows, and orchestration frameworks to support efficient model development and deployment.
  • Design and implement high-performance model serving infrastructure using industry-standard serving frameworks while optimizing inference for low latency and high throughput.
  • Build reliable deployment strategies including A/B testing, canary releases, rollback mechanisms, and production validation processes.
  • Create robust monitoring, logging, alerting, and observability solutions to ensure model reliability, performance, and operational excellence.
  • Optimize infrastructure utilization by improving GPU efficiency, enabling autoscaling, and managing cloud resources effectively.
  • Design and maintain feature stores, scalable data pipelines, and storage architectures capable of supporting large-scale training and inference workloads.
  • Collaborate with engineering teams to continuously improve platform scalability, security, governance, and operational best practices.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now