Senior Inference Engineer
New
V
vCluster LabsCloud Infrastructure
Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe and we have a remote-first work culture.Full-TimeSenior
SalaryUnited States Salary: $190K – $230K; Germany Salary: €190K – €230K
Apply NowOpens the employer's application page
Job Details
- Required Skills
- DockerPythonKubernetesPyTorchGo
Requirements
- Production-level LLM serving experience.
- Hands-on experience deploying with vLLM, SGLang, or TensorRT-LLM.
- Deep technical expertise in inference optimization (quantization, batching, caching, routing).
- Strong proficiency in Python or Golang for building production infrastructure.
- Ability to explain complex technical concepts to diverse stakeholders.
- Understanding of containerized environments (Docker, Kubernetes).
- Experience with ML frameworks such as PyTorch and Transformers.
- Knowledge of GPU stacks including CUDA, NCCL, and drivers.
- Familiarity with model architectures and fine-tuning approaches.
Responsibilities
- Deploy LLMs into production on GPU infrastructure.
- Own the full pipeline from customer query to served response.
- Operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
- Optimize latency and cost through quantization, batching, caching, and routing.
- Build infrastructure using Python or Golang.
- Partner with the CTO and Product team to define and drive the inference platform roadmap.
View Full Description & ApplyYou'll be redirected to the employer's site