Senior ML Engineer

New
C
Cast AICloud Infrastructure
Enjoy a flexible, remote-first global environment.Full-TimeSenior
SalaryCompetitive salary (depending on the level of experience).
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
PythonKubernetesMachine LearningDistributed Systems

Requirements

  • 5+ years of experience building real ML systems.
  • Portfolio demonstrating depth in inference or training infrastructure.
  • Strong Python programming skills suitable for production services.
  • Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM.
  • Understanding of inference engine performance on specific GPU architectures.
  • Fluency with quantization tradeoffs and measuring quality regressions.
  • Comfort with distributed systems, including collective communication and sharding.
  • Ability to diagnose failure modes in multi-GPU and multi-node setups.
  • Bias toward measurement and identifying performance wins vs. benchmarks.
  • Strong self-direction and ability to lead technical strategy.

Responsibilities

  • Maximize throughput using techniques like continuous batching, speculative decoding, and kernel-level tuning.
  • Reduce latency by profiling and addressing bottlenecks in TTFT and TPOT.
  • Optimize KV cache utilization through paged attention, prefix caching, and cache reuse strategies.
  • Perform quantization (INT8, INT4, FP8) while ensuring quality on real-world workloads.
  • Improve cold start times and memory footprint through efficient weight loading and accounting.
  • Scale inference across distributed topologies with network-aware placement and checkpointing.
  • Define the technical direction and benchmarks for LLM inference optimization.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive salary (depending on the level of experience).
Apply Now