AI Platform Engineer

A
AHEADAI/ML infrastructure
Listing location: India; Workplace type: Remote; Structured job location: IndiaFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
4+ years in platform architecture or solutions architecture, with 2+ years focused on AI/ML workloads.
Required Skills
PythonKubernetesPyTorchTensorflowTerraformAnsibleMLOps

Requirements

  • Have 4+ years of experience in platform architecture or solutions architecture, including 2+ years focused on AI/ML workloads.
  • Demonstrate strong proficiency in Python.
  • Have experience with TensorFlow or PyTorch.
  • Have hands-on experience with Kubernetes and container orchestration.
  • Be familiar with Run:ai or similar GPU scheduling platforms.
  • Have expertise in Terraform and Ansible for infrastructure automation.
  • Have experience using Jupyter Notebooks for ML development.
  • Know NVIDIA Enterprise Suite components, including CUDA, NeMo Framework, Triton, and GPU drivers.
  • Understand MLOps principles and tools such as MLflow and Kubeflow.
  • Have a background deploying and scaling AI workloads in cloud or hybrid environments.
  • Have experience with high-performance computing (HPC) environments.
  • Be familiar with distributed training and model optimization techniques.
  • Hold a Kubernetes or cloud platform certification (AWS, Azure, or GCP).

Responsibilities

  • Architect and manage Kubernetes clusters tailored to AI/ML workloads.
  • Implement Run:ai and operators for GPU resource orchestration and workload scheduling.
  • Develop and maintain Python-based automation scripts and ML pipelines.
  • Automate infrastructure provisioning with Terraform and configuration management with Ansible.
  • Create and manage Jupyter Notebooks for experimentation and collaboration.
  • Integrate and optimize NVIDIA Enterprise Suite components, including CUDA, NeMo Framework, Triton, TensorRT, and GPU drivers.
  • Establish and maintain MLOps practices for model lifecycle management, CI/CD, and monitoring.
  • Collaborate with data scientists and platform engineers to support resource utilization and scalability across cloud and hybrid environments.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now