AI Research Engineer (Model Compression & Quantization)
T
TetherFintech / AI
100% Remote WorldwideFull-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Required Skills
- Machine LearningNLP
Requirements
- Degree in Computer Science or related field; PhD in NLP or Machine Learning preferred.
- Proven track record in AI R&D with publications in A* conferences.
- Expertise in Metal Shading Language (MSL) and writing custom compute shaders from scratch.
- Experience in low-level kernel optimizations and inference optimization on mobile devices.
- Deep understanding of modern model serving architectures and inference optimization techniques.
- Strong expertise in writing GPU kernels for mobile devices.
- Proficiency in designing robust evaluation frameworks for latency and memory constraints.
- Knowledge of distributed inference systems, including Tensor, Pipeline, and Expert Parallelism.
- Understanding of Diffusion Models, Vision Transformers, Pruning, Quantization, Flash attention, KV Cache, and Speculative Decoding.
- Excellent English communication skills.
Responsibilities
- Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage.
- Build, run, and monitor controlled inference tests in both simulated and live production environments.
- Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges.
- Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring processing and memory metrics.
- Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines.
View Full Description & ApplyYou'll be redirected to the employer's site