AI Engineer - Model Performance

New
F
FathomAI/Software
We embrace being fully remote.Full-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
Python

Requirements

  • Deep experience with LLM serving frameworks (vLLM, SGLang, TensorRT-LLM) including attention backends and scheduling.
  • Hands-on experience with weight vs activation quantization and scaling techniques.
  • Production fine-tuning experience (LoRA/QLoRA SFT) and familiarity with training frameworks (ms-swift, Axolotl, torchtune).
  • Strong Python programming skills for building infrastructure, benchmarks, and pipelines.
  • Expertise in GPU profiling, performance analysis, and identifying compute vs. memory bottlenecks.
  • Understanding of data formatting and training hyperparameters for model fine-tuning.
  • Ability to build internal developer tooling to improve team velocity.
  • Strong communication skills for documenting technical tradeoffs and decisions.

Responsibilities

  • Optimize model inference speed, cost, and reliability for production systems serving millions of meetings.
  • Benchmark and implement quantization strategies including FP8 and static vs. dynamic quantization.
  • Evaluate and configure serving frameworks such as vLLM and SGLang.
  • Develop repeatable fine-tuning infrastructure for tasks like classification and adapter training.
  • Model GPU infrastructure costs and make strategic hardware selection decisions.
  • Debug production inference and quality regressions to ensure system stability.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now