Machine Learning Engineer — Training Optimization
New
J
JobgetherMachine Learning
IndiaFull-Time
SalaryCompetitive compensation package with meaningful equity opportunities.
Apply NowOpens the employer's application page
Job Details
- Required Skills
- Machine LearningPyTorchDistributed Systems
Requirements
- Strong experience training large-scale neural networks, including large language models or similarly complex architectures.
- Hands-on experience with machine learning training optimization, beyond simply using existing models.
- Strong understanding of backpropagation, optimization algorithms, training dynamics, and model convergence behavior.
- Experience with distributed machine learning training systems and large-scale computing environments.
- Proficiency with PyTorch and modern machine learning development workflows.
- Ability to work close to hardware constraints, including GPU performance, memory limitations, and networking considerations.
- Strong programming skills with the ability to transform research ideas into reliable production code.
- Experience with multi-node and multi-GPU training environments is highly preferred.
- Familiarity with frameworks and technologies such as DeepSpeed, FSDP, Megatron, or custom training stacks is a plus.
- Experience optimizing workloads on NVIDIA or AMD GPU platforms is beneficial.
Responsibilities
- Optimize large-scale model training pipelines to improve throughput, convergence, stability, and overall computational efficiency.
- Improve distributed training approaches, including data parallelism, model parallelism, and pipeline parallelism strategies.
- Tune key training components such as optimizers, learning rate schedulers, batch sizes, and numerical precision methods including bf16, fp16, and fp8.
- Identify and resolve performance bottlenecks through profiling, system analysis, and infrastructure-level improvements.
- Collaborate closely with research teams to develop architecture-aware training strategies and improve model performance.
- Build and maintain reliable training infrastructure, including checkpointing systems, fault tolerance mechanisms, and reproducible workflows.
- Evaluate and integrate advanced training techniques such as gradient checkpointing, ZeRO, FSDP, and custom optimization solutions.
- Define, monitor, and improve training performance metrics to continuously enhance efficiency.
- Translate research concepts into production-ready systems and scalable engineering solutions.
View Full Description & ApplyYou'll be redirected to the employer's site