Senior ML Engineer, LLM Inference Optimization

New
C
Cast AICloud Infrastructure
Cast AI operates across 34 countries spanning Europe, North America, Latin America, and APACFull-TimeSenior
SalaryCompetitive salary (depending on the level of experience).
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
PythonKubernetesPyTorchDistributed Systems

Requirements

  • 5+ years of experience building real ML systems.
  • Demonstrated depth in inference or training infrastructure.
  • Strong proficiency in production-grade Python.
  • Hands-on experience with at least one of: vLLM, SGLang, or TensorRT-LLM.
  • Understanding of inference engine performance dynamics on GPU hardware.
  • Fluency in quantization tradeoffs and quality measurement.
  • Experience with distributed systems, collective communication, and sharding strategies.
  • Strong focus on measurement and benchmarking instrumentation.
  • Ability to work with self-direction in a wide-mandate role.

Responsibilities

  • Push throughput via continuous batching, speculative decoding, chunked prefill, and kernel-level tuning.
  • Cut latency by profiling and addressing bottlenecks in compute, memory bandwidth, scheduling, and networking.
  • Optimize KV cache usage through paged attention, prefix caching, and eviction policies.
  • Implement quantization (INT8, INT4, FP8) while ensuring quality across real workloads.
  • Minimize cold starts and memory footprints through efficient initialization and weight loading.
  • Scale distributed inference topologies and network-aware placement strategies.
  • Define the technical roadmap for benchmarking, model adoption, and internal tooling development.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive salary (depending on the level of experience).
Apply Now