Engineering Lead, Inference Optimization
New
V
VeniceArtificial Intelligence
Remote- US onlyFull-TimeLead
Salary270,000 - 330,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years in performance optimization or HPC; 5+ years experience leading engineering teams
- Required Skills
- PythonFGPA ArchitectureGoRust
Requirements
- 8+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge.
- 5+ years experience leading engineering teams.
- Proficiency in Python, Rust, or Go.
- Hands-on experience with at least one production LLM inference engine (e.g. vLLM, SGLang) running at high volume.
- Demonstrated experience with LLM inference optimization techniques: continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile.
- Fluency with quantization tradeoffs, both qualitative and quantitative.
- Experience with distributed inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments.
- Fluency with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler).
- C++/CUDA experience (Bonus).
- Hands-on kernel development experience (Bonus).
- Experience with diffusion/image model inference optimization, custom Triton kernels, or contributions to open-source inference frameworks (Bonus).
Responsibilities
- Own Venice’s technical strategy for inference performance.
- Recruit and lead the Inference Optimization Team at Venice.
- Optimize Venice's GPU infrastructure across a range of architectures (e.g. H200s, B300s).
- Improve latency, throughput, and cost per token for LLM inference workloads.
- Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang).
- Work with inference routing system to optimize multivariate inference load-balancing algorithms.
- Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels) and compilation stack improvements.
- Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability.
View Full Description & ApplyYou'll be redirected to the employer's site