Senior ML Engineer
New
C
Cast AICloud Infrastructure
Enjoy a flexible, remote-first global environment.Full-TimeSenior
SalaryCompetitive salary (depending on the level of experience).
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonKubernetesMachine LearningDistributed Systems
Requirements
- 5+ years of experience building real ML systems.
- Portfolio demonstrating depth in inference or training infrastructure.
- Strong Python programming skills suitable for production services.
- Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM.
- Understanding of inference engine performance on specific GPU architectures.
- Fluency with quantization tradeoffs and measuring quality regressions.
- Comfort with distributed systems, including collective communication and sharding.
- Ability to diagnose failure modes in multi-GPU and multi-node setups.
- Bias toward measurement and identifying performance wins vs. benchmarks.
- Strong self-direction and ability to lead technical strategy.
Responsibilities
- Maximize throughput using techniques like continuous batching, speculative decoding, and kernel-level tuning.
- Reduce latency by profiling and addressing bottlenecks in TTFT and TPOT.
- Optimize KV cache utilization through paged attention, prefix caching, and cache reuse strategies.
- Perform quantization (INT8, INT4, FP8) while ensuring quality on real-world workloads.
- Improve cold start times and memory footprint through efficient weight loading and accounting.
- Scale inference across distributed topologies with network-aware placement and checkpointing.
- Define the technical direction and benchmarks for LLM inference optimization.
View Full Description & ApplyYou'll be redirected to the employer's site