Senior Inference Engineer

New
V
vCluster LabsCloud Infrastructure
Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe and we have a remote-first work culture.Full-TimeSenior
SalaryUnited States Salary: $190K – $230K; Germany Salary: €190K – €230K
Apply NowOpens the employer's application page

Job Details

Required Skills
DockerPythonKubernetesPyTorchGo

Requirements

  • Production-level LLM serving experience.
  • Hands-on experience deploying with vLLM, SGLang, or TensorRT-LLM.
  • Deep technical expertise in inference optimization (quantization, batching, caching, routing).
  • Strong proficiency in Python or Golang for building production infrastructure.
  • Ability to explain complex technical concepts to diverse stakeholders.
  • Understanding of containerized environments (Docker, Kubernetes).
  • Experience with ML frameworks such as PyTorch and Transformers.
  • Knowledge of GPU stacks including CUDA, NCCL, and drivers.
  • Familiarity with model architectures and fine-tuning approaches.

Responsibilities

  • Deploy LLMs into production on GPU infrastructure.
  • Own the full pipeline from customer query to served response.
  • Operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Optimize latency and cost through quantization, batching, caching, and routing.
  • Build infrastructure using Python or Golang.
  • Partner with the CTO and Product team to define and drive the inference platform roadmap.
View Full Description & ApplyYou'll be redirected to the employer's site
United States Salary: $190K – $230K; Germany Salary: €190K – €230K
Apply Now