Senior AI ML Operations Engineer

New
J
JobgetherAI/ML Operations
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSPythonSQLKubernetesMLFlowCI/CDTerraformDatabricksMLOps

Requirements

  • Proven experience building and maintaining scalable AI/ML Operations platforms in enterprise environments.
  • Experience working with large-scale structured and unstructured data platforms.
  • In-depth experience with Databricks Lakehouse ecosystem and MLflow.
  • Knowledge of Mosaic AI Agent Framework, Unity Catalog, Vector Search, and Knowledge Graph technologies.
  • Experience integrating pipelines with LangChain and LangGraph.
  • Strong programming skills in Python and SQL.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
  • Experience with Kubernetes, CI/CD, and Infrastructure as Code (preferably Terraform).
  • Experience with model deployment, production monitoring, and data drift detection.
  • Understanding of foundation models, fine-tuning, RAG, and vector databases.
  • Understanding of cloud resource management and infrastructure optimization.
  • Experience supporting highly available and scalable enterprise SaaS or technology platforms.

Responsibilities

  • Design, develop, and maintain stable, scalable, reliable, and secure AI/ML Operations platforms and production pipelines.
  • Package, deploy, and manage AI/ML services in production ensuring reproducible and maintainable deployments.
  • Build and optimize CI/CD pipelines to automate model deployment and accelerate production transitions.
  • Provision and manage infrastructure for training and inference using Docker, Kubernetes, and Infrastructure as Code.
  • Implement monitoring and observability for model performance, data drift, and production metrics.
  • Automate model retraining and data workflows to maintain operational consistency.
  • Manage foundation model deployments, fine-tuning, and RAG architectures involving vector databases.
  • Optimize infrastructure resources and GPU/CPU utilization for cost-efficiency and low-latency inference.
  • Partner with data scientists and engineers to bridge development and production operations.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now