Senior AI ML Operations Engineer

New
J
JobgetherAI/ML Operations
IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
DockerPythonSQLKubernetesMLFlowCI/CDTerraformDatabricksLangChain

Requirements

  • Proven experience working with enterprise SaaS solutions requiring high availability, scalability, reliability, and performance.
  • Strong experience building and maintaining AI/ML Operations platforms and production workflows at scale.
  • In-depth experience handling large-scale structured and unstructured data.
  • Hands-on expertise with the Databricks Lakehouse ecosystem, including MLflow, Unity Catalog, and Vector Search.
  • Experience with the Mosaic AI Agent Framework, knowledge graphs, and modern AI/ML platform architectures.
  • Strong knowledge of AI/ML frameworks such as LangChain and LangGraph.
  • Practical experience deploying and operating AI/ML workloads on at least one major cloud platform (AWS, Azure, or GCP).
  • Strong programming skills in Python and SQL.
  • Experience with modern software engineering practices including Kubernetes, Docker, CI/CD, and infrastructure as code.
  • Experience with infrastructure-as-code tools such as Terraform.
  • Strong understanding of monitoring, alerting, reliability engineering, and production support for AI/ML platforms.

Responsibilities

  • Design, develop, and maintain stable, scalable, secure, and reliable AI/ML Operations platforms and production pipelines.
  • Package, deploy, and manage AI/ML services and models in production environments, ensuring reproducibility, reliability, and interpretability.
  • Design and implement automated CI/CD pipelines to streamline model deployment and operational workflows.
  • Provision, configure, and optimize infrastructure for AI/ML training and inference using Docker, Kubernetes, serverless technologies, and infrastructure-as-code practices.
  • Implement monitoring and observability for model performance, data drift, latency, reliability, and platform health.
  • Manage deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation solutions.
  • Optimize GPU and CPU resources to control cloud costs while maintaining high-performance inference.
  • Establish effective version control and governance for data, code, models, and AI/ML assets.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now