Intelligent Infrastructure Engineer

J
JobgetherAI Infrastructure
Based in the United StatesFull-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
PythonKubernetesC++GoLinuxDistributed Systems

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical discipline.
  • 6+ years of experience in infrastructure, platform engineering, HPC engineering, or related fields.
  • Hands-on experience operating GPU clusters or large-scale machine learning infrastructure.
  • Strong proficiency in Python and experience with a systems programming language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and communication frameworks.
  • Experience with Kubernetes, Slurm, Ray, or comparable workload scheduling systems.
  • Strong knowledge of Linux internals, networking concepts, and high-performance storage technologies.
  • Experience working with major cloud providers and their machine learning infrastructure offerings.
  • Strong software engineering skills, including testing, automation, CI/CD, and collaborative development practices.

Responsibilities

  • Design, build, and operate infrastructure platforms supporting large-scale AI model training and inference workloads.
  • Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
  • Improve platform reliability, performance, scalability, and operational efficiency across AI infrastructure.
  • Develop software solutions and automation tools using Python and systems programming languages such as Go or C++.
  • Work with distributed training frameworks, accelerator architectures, and high-performance computing environments.
  • Configure and optimize Kubernetes, Slurm, Ray, or similar orchestration platforms for ML workloads.
  • Troubleshoot and improve Linux-based infrastructure, networking systems, and high-performance storage solutions.
  • Partner with machine learning engineers, researchers, and infrastructure teams to improve developer workflows and platform capabilities.
  • Implement strong software engineering practices, including testing, CI/CD processes, and code reviews.
  • Contribute to cost optimization initiatives and operational improvements for AI workloads.
View Full Description & ApplyYou'll be redirected to the employer's site
100,000 - 150,000 USD per year
Apply Now