Senior AI ML Operations Engineer
New
J
JobgetherAI/ML Operations
IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- DockerPythonSQLKubernetesMLFlowCI/CDTerraformDatabricksLangChain
Requirements
- Proven experience working with enterprise SaaS solutions requiring high availability, scalability, reliability, and performance.
- Strong experience building and maintaining AI/ML Operations platforms and production workflows at scale.
- In-depth experience handling large-scale structured and unstructured data.
- Hands-on expertise with the Databricks Lakehouse ecosystem, including MLflow, Unity Catalog, and Vector Search.
- Experience with the Mosaic AI Agent Framework, knowledge graphs, and modern AI/ML platform architectures.
- Strong knowledge of AI/ML frameworks such as LangChain and LangGraph.
- Practical experience deploying and operating AI/ML workloads on at least one major cloud platform (AWS, Azure, or GCP).
- Strong programming skills in Python and SQL.
- Experience with modern software engineering practices including Kubernetes, Docker, CI/CD, and infrastructure as code.
- Experience with infrastructure-as-code tools such as Terraform.
- Strong understanding of monitoring, alerting, reliability engineering, and production support for AI/ML platforms.
Responsibilities
- Design, develop, and maintain stable, scalable, secure, and reliable AI/ML Operations platforms and production pipelines.
- Package, deploy, and manage AI/ML services and models in production environments, ensuring reproducibility, reliability, and interpretability.
- Design and implement automated CI/CD pipelines to streamline model deployment and operational workflows.
- Provision, configure, and optimize infrastructure for AI/ML training and inference using Docker, Kubernetes, serverless technologies, and infrastructure-as-code practices.
- Implement monitoring and observability for model performance, data drift, latency, reliability, and platform health.
- Manage deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation solutions.
- Optimize GPU and CPU resources to control cloud costs while maintaining high-performance inference.
- Establish effective version control and governance for data, code, models, and AI/ML assets.
View Full Description & ApplyYou'll be redirected to the employer's site