Senior ML Engineer, LLM Inference Optimization
New
C
Cast AICloud Infrastructure
Cast AI operates across 34 countries spanning Europe, North America, Latin America, and APACFull-TimeSenior
SalaryCompetitive salary (depending on the level of experience).
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonKubernetesPyTorchDistributed Systems
Requirements
- 5+ years of experience building real ML systems.
- Demonstrated depth in inference or training infrastructure.
- Strong proficiency in production-grade Python.
- Hands-on experience with at least one of: vLLM, SGLang, or TensorRT-LLM.
- Understanding of inference engine performance dynamics on GPU hardware.
- Fluency in quantization tradeoffs and quality measurement.
- Experience with distributed systems, collective communication, and sharding strategies.
- Strong focus on measurement and benchmarking instrumentation.
- Ability to work with self-direction in a wide-mandate role.
Responsibilities
- Push throughput via continuous batching, speculative decoding, chunked prefill, and kernel-level tuning.
- Cut latency by profiling and addressing bottlenecks in compute, memory bandwidth, scheduling, and networking.
- Optimize KV cache usage through paged attention, prefix caching, and eviction policies.
- Implement quantization (INT8, INT4, FP8) while ensuring quality across real workloads.
- Minimize cold starts and memory footprints through efficient initialization and weight loading.
- Scale distributed inference topologies and network-aware placement strategies.
- Define the technical roadmap for benchmarking, model adoption, and internal tooling development.
View Full Description & ApplyYou'll be redirected to the employer's site