- Optimize model inference speed, cost, and reliability for production systems serving millions of meetings.
- Benchmark and implement quantization strategies including FP8 and static vs. dynamic quantization.
- Evaluate and configure serving frameworks such as vLLM and SGLang.
- Develop repeatable fine-tuning infrastructure for tasks like classification and adapter training.
- Model GPU infrastructure costs and make strategic hardware selection decisions.
- Debug production inference and quality regressions to ensure system stability.
Python