Senior ML Systems Engineer, Inference
R
RunpodAI Infrastructure
Remote - USAFull-TimeSenior
Salary$150,000 - $220,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of professional system engineering experience
- Required Skills
- Python
Requirements
- 5+ years of professional system engineering experience.
- Deep, hands-on experience with vLLM, SGLang, or comparable serving engines in production or at benchmark scale.
- Strong software engineering skills in Python.
- Comfortable working in large, performance-critical codebases.
- Solid understanding of LLM inference performance drivers (batching, memory, parallelism, and latency/throughput trade-offs).
- Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.
- Rigor in benchmarking and performance analysis.
- Experience with GPU profiling tools.
- Ability to explain technical results clearly in writing and turn them into actionable decisions.
- Preferred: Experience writing or tuning GPU kernels in CUDA or Triton.
- Preferred: Contributions to inference or ML systems projects.
- Preferred: Experience with multi-node GPU systems and high-speed networking.
- Preferred: Experience at a company where inference cost and latency were core business metrics.
Responsibilities
- Define and build tooling for measuring inference performance, including throughput, time to first token, inter-token latency, and cost per token.
- Profile and diagnose performance issues across the entire serving stack, from scheduling to kernels and interconnect.
- Improve serving efficiency for state-of-the-art models on single-node and multi-node GPU deployments.
- Develop production-ready runtimes, configurations, and defaults to optimize customer inference.
- Collaborate with product and infrastructure teams to shape Runpod's inference offerings.
- Engage with the inference open-source community to evaluate, build, or contribute performance optimizations.
- Trace and implement low-level fixes in the serving engine when configuration tuning is insufficient.
View Full Description & ApplyYou'll be redirected to the employer's site