Senior AI ML Operations Engineer
New
J
JobgetherAI/ML Operations
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSPythonSQLKubernetesMLFlowCI/CDTerraformDatabricksMLOps
Requirements
- Proven experience building and maintaining scalable AI/ML Operations platforms in enterprise environments.
- Experience working with large-scale structured and unstructured data platforms.
- In-depth experience with Databricks Lakehouse ecosystem and MLflow.
- Knowledge of Mosaic AI Agent Framework, Unity Catalog, Vector Search, and Knowledge Graph technologies.
- Experience integrating pipelines with LangChain and LangGraph.
- Strong programming skills in Python and SQL.
- Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
- Experience with Kubernetes, CI/CD, and Infrastructure as Code (preferably Terraform).
- Experience with model deployment, production monitoring, and data drift detection.
- Understanding of foundation models, fine-tuning, RAG, and vector databases.
- Understanding of cloud resource management and infrastructure optimization.
- Experience supporting highly available and scalable enterprise SaaS or technology platforms.
Responsibilities
- Design, develop, and maintain stable, scalable, reliable, and secure AI/ML Operations platforms and production pipelines.
- Package, deploy, and manage AI/ML services in production ensuring reproducible and maintainable deployments.
- Build and optimize CI/CD pipelines to automate model deployment and accelerate production transitions.
- Provision and manage infrastructure for training and inference using Docker, Kubernetes, and Infrastructure as Code.
- Implement monitoring and observability for model performance, data drift, and production metrics.
- Automate model retraining and data workflows to maintain operational consistency.
- Manage foundation model deployments, fine-tuning, and RAG architectures involving vector databases.
- Optimize infrastructure resources and GPU/CPU utilization for cost-efficiency and low-latency inference.
- Partner with data scientists and engineers to bridge development and production operations.
View Full Description & ApplyYou'll be redirected to the employer's site