AI Engineer - Model Performance
New
F
FathomAI/Software
We embrace being fully remote.Full-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- Python
Requirements
- Deep experience with LLM serving frameworks (vLLM, SGLang, TensorRT-LLM) including attention backends and scheduling.
- Hands-on experience with weight vs activation quantization and scaling techniques.
- Production fine-tuning experience (LoRA/QLoRA SFT) and familiarity with training frameworks (ms-swift, Axolotl, torchtune).
- Strong Python programming skills for building infrastructure, benchmarks, and pipelines.
- Expertise in GPU profiling, performance analysis, and identifying compute vs. memory bottlenecks.
- Understanding of data formatting and training hyperparameters for model fine-tuning.
- Ability to build internal developer tooling to improve team velocity.
- Strong communication skills for documenting technical tradeoffs and decisions.
Responsibilities
- Optimize model inference speed, cost, and reliability for production systems serving millions of meetings.
- Benchmark and implement quantization strategies including FP8 and static vs. dynamic quantization.
- Evaluate and configure serving frameworks such as vLLM and SGLang.
- Develop repeatable fine-tuning infrastructure for tasks like classification and adapter training.
- Model GPU infrastructure costs and make strategic hardware selection decisions.
- Debug production inference and quality regressions to ensure system stability.
View Full Description & ApplyYou'll be redirected to the employer's site