MLOps Engineer
New
J
JobgetherMachine Learning
Germany; Fully remote work within EuropeFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- At least 5 years
- Required Skills
- PythonKubeflowKubernetesMLFlowPyTorchGoTensorflowCI/CDTerraformMLOps
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
- At least 5 years of experience in MLOps, DevOps, or related engineering roles supporting production machine learning environments.
- Proven experience designing and building MLOps infrastructure from the ground up using platforms such as MLflow, Weights & Biases, Kubeflow, or similar.
- Strong hands-on experience with machine learning frameworks including PyTorch and TensorFlow.
- Experience with model serving technologies such as TorchServe, TensorFlow Serving, Triton, or KServe.
- Solid experience developing and managing scalable data pipelines and Kubernetes environments.
- Experience with cloud infrastructure (AWS, GCP, or Azure) and Infrastructure as Code solutions including Terraform, Helm, or GitOps.
- Strong programming skills in Python, Bash, and Go, with a focus on maintainable, scalable, and production-quality software.
- Knowledge of AI system security, model governance, compliance, monitoring, and observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
Responsibilities
- Develop, automate, and maintain scalable machine learning pipelines, CI/CD workflows, and orchestration frameworks to support efficient model development and deployment.
- Design and implement high-performance model serving infrastructure using industry-standard serving frameworks while optimizing inference for low latency and high throughput.
- Build reliable deployment strategies including A/B testing, canary releases, rollback mechanisms, and production validation processes.
- Create robust monitoring, logging, alerting, and observability solutions to ensure model reliability, performance, and operational excellence.
- Optimize infrastructure utilization by improving GPU efficiency, enabling autoscaling, and managing cloud resources effectively.
- Design and maintain feature stores, scalable data pipelines, and storage architectures capable of supporting large-scale training and inference workloads.
- Collaborate with engineering teams to continuously improve platform scalability, security, governance, and operational best practices.
View Full Description & ApplyYou'll be redirected to the employer's site