Senior Machine Learning Engineer, LLM Inference Optimization
New
J
JobgetherArtificial Intelligence
Based in United KingdomFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonPyTorchLLM
Requirements
- Strong software engineering skills in Python and PyTorch.
- Hands-on experience deploying, operating, or optimizing LLM, VLM, or high-throughput transformer inference systems.
- Practical experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, or KServe.
- Strong understanding of transformer inference bottlenecks including KV cache, attention mechanisms, memory bandwidth, and parallelism.
- Ability to reason quantitatively about latency, throughput, model quality, resource utilization, and cost trade-offs.
- Experience diagnosing complex performance problems and translating findings into production improvements.
- Strong communication skills and ability to collaborate with cross-functional teams.
- Experience with quantization techniques (FP8, INT8, AWQ, GPTQ, etc.) is a plus.
- Familiarity with inference acceleration approaches like speculative decoding or multi-token prediction is advantageous.
- Ability to work independently and operate effectively in a fast-moving technical environment.
Responsibilities
- Own optimization initiatives for specific model families, customer endpoints, and inference serving backends.
- Evaluate inference engines and recommend practical serving configurations based on workload requirements.
- Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token.
- Deploy, configure, benchmark, and extend modern inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or NVIDIA Dynamo.
- Build and productionize model-compression workflows including quantization, distillation, and low-bit serving.
- Develop reproducible benchmark harnesses for latency, throughput, and GPU resource usage.
- Partner with GPU kernel and platform engineers to identify and resolve performance bottlenecks.
View Full Description & ApplyYou'll be redirected to the employer's site