Senior Machine Learning Engineer, LLM Inference Optimization

New
J
JobgetherArtificial Intelligence
Based in United KingdomFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonPyTorchLLM

Requirements

  • Strong software engineering skills in Python and PyTorch.
  • Hands-on experience deploying, operating, or optimizing LLM, VLM, or high-throughput transformer inference systems.
  • Practical experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, or KServe.
  • Strong understanding of transformer inference bottlenecks including KV cache, attention mechanisms, memory bandwidth, and parallelism.
  • Ability to reason quantitatively about latency, throughput, model quality, resource utilization, and cost trade-offs.
  • Experience diagnosing complex performance problems and translating findings into production improvements.
  • Strong communication skills and ability to collaborate with cross-functional teams.
  • Experience with quantization techniques (FP8, INT8, AWQ, GPTQ, etc.) is a plus.
  • Familiarity with inference acceleration approaches like speculative decoding or multi-token prediction is advantageous.
  • Ability to work independently and operate effectively in a fast-moving technical environment.

Responsibilities

  • Own optimization initiatives for specific model families, customer endpoints, and inference serving backends.
  • Evaluate inference engines and recommend practical serving configurations based on workload requirements.
  • Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token.
  • Deploy, configure, benchmark, and extend modern inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or NVIDIA Dynamo.
  • Build and productionize model-compression workflows including quantization, distillation, and low-bit serving.
  • Develop reproducible benchmark harnesses for latency, throughput, and GPU resource usage.
  • Partner with GPU kernel and platform engineers to identify and resolve performance bottlenecks.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now