Engineering Lead, Inference Optimization

New
V
VeniceArtificial Intelligence
Remote- US onlyFull-TimeLead
Salary270,000 - 330,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
8+ years in performance optimization or HPC; 5+ years experience leading engineering teams
Required Skills
PythonFGPA ArchitectureGoRust

Requirements

  • 8+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge.
  • 5+ years experience leading engineering teams.
  • Proficiency in Python, Rust, or Go.
  • Hands-on experience with at least one production LLM inference engine (e.g. vLLM, SGLang) running at high volume.
  • Demonstrated experience with LLM inference optimization techniques: continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile.
  • Fluency with quantization tradeoffs, both qualitative and quantitative.
  • Experience with distributed inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments.
  • Fluency with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler).
  • C++/CUDA experience (Bonus).
  • Hands-on kernel development experience (Bonus).
  • Experience with diffusion/image model inference optimization, custom Triton kernels, or contributions to open-source inference frameworks (Bonus).

Responsibilities

  • Own Venice’s technical strategy for inference performance.
  • Recruit and lead the Inference Optimization Team at Venice.
  • Optimize Venice's GPU infrastructure across a range of architectures (e.g. H200s, B300s).
  • Improve latency, throughput, and cost per token for LLM inference workloads.
  • Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang).
  • Work with inference routing system to optimize multivariate inference load-balancing algorithms.
  • Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels) and compilation stack improvements.
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability.
View Full Description & ApplyYou'll be redirected to the employer's site
270,000 - 330,000 USD per year
Apply Now