Machine Learning Engineer — Training Optimization

New
J
JobgetherMachine Learning
IndiaFull-Time
SalaryCompetitive compensation package with meaningful equity opportunities.
Apply NowOpens the employer's application page

Job Details

Required Skills
Machine LearningPyTorchDistributed Systems

Requirements

  • Strong experience training large-scale neural networks, including large language models or similarly complex architectures.
  • Hands-on experience with machine learning training optimization, beyond simply using existing models.
  • Strong understanding of backpropagation, optimization algorithms, training dynamics, and model convergence behavior.
  • Experience with distributed machine learning training systems and large-scale computing environments.
  • Proficiency with PyTorch and modern machine learning development workflows.
  • Ability to work close to hardware constraints, including GPU performance, memory limitations, and networking considerations.
  • Strong programming skills with the ability to transform research ideas into reliable production code.
  • Experience with multi-node and multi-GPU training environments is highly preferred.
  • Familiarity with frameworks and technologies such as DeepSpeed, FSDP, Megatron, or custom training stacks is a plus.
  • Experience optimizing workloads on NVIDIA or AMD GPU platforms is beneficial.

Responsibilities

  • Optimize large-scale model training pipelines to improve throughput, convergence, stability, and overall computational efficiency.
  • Improve distributed training approaches, including data parallelism, model parallelism, and pipeline parallelism strategies.
  • Tune key training components such as optimizers, learning rate schedulers, batch sizes, and numerical precision methods including bf16, fp16, and fp8.
  • Identify and resolve performance bottlenecks through profiling, system analysis, and infrastructure-level improvements.
  • Collaborate closely with research teams to develop architecture-aware training strategies and improve model performance.
  • Build and maintain reliable training infrastructure, including checkpointing systems, fault tolerance mechanisms, and reproducible workflows.
  • Evaluate and integrate advanced training techniques such as gradient checkpointing, ZeRO, FSDP, and custom optimization solutions.
  • Define, monitor, and improve training performance metrics to continuously enhance efficiency.
  • Translate research concepts into production-ready systems and scalable engineering solutions.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive compensation package with meaningful equity opportunities.
Apply Now