Intelligent Infrastructure Engineer
J
JobgetherAI Infrastructure
Based in the United StatesFull-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- PythonKubernetesC++GoLinuxDistributed Systems
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical discipline.
- 6+ years of experience in infrastructure, platform engineering, HPC engineering, or related fields.
- Hands-on experience operating GPU clusters or large-scale machine learning infrastructure.
- Strong proficiency in Python and experience with a systems programming language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, and communication frameworks.
- Experience with Kubernetes, Slurm, Ray, or comparable workload scheduling systems.
- Strong knowledge of Linux internals, networking concepts, and high-performance storage technologies.
- Experience working with major cloud providers and their machine learning infrastructure offerings.
- Strong software engineering skills, including testing, automation, CI/CD, and collaborative development practices.
Responsibilities
- Design, build, and operate infrastructure platforms supporting large-scale AI model training and inference workloads.
- Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
- Improve platform reliability, performance, scalability, and operational efficiency across AI infrastructure.
- Develop software solutions and automation tools using Python and systems programming languages such as Go or C++.
- Work with distributed training frameworks, accelerator architectures, and high-performance computing environments.
- Configure and optimize Kubernetes, Slurm, Ray, or similar orchestration platforms for ML workloads.
- Troubleshoot and improve Linux-based infrastructure, networking systems, and high-performance storage solutions.
- Partner with machine learning engineers, researchers, and infrastructure teams to improve developer workflows and platform capabilities.
- Implement strong software engineering practices, including testing, CI/CD processes, and code reviews.
- Contribute to cost optimization initiatives and operational improvements for AI workloads.
View Full Description & ApplyYou'll be redirected to the employer's site