AI Research Engineer (Model Compression & Quantization)

T
TetherFintech / AI
100% Remote WorldwideFull-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English
Required Skills
Machine LearningNLP

Requirements

  • Degree in Computer Science or related field; PhD in NLP or Machine Learning preferred.
  • Proven track record in AI R&D with publications in A* conferences.
  • Expertise in Metal Shading Language (MSL) and writing custom compute shaders from scratch.
  • Experience in low-level kernel optimizations and inference optimization on mobile devices.
  • Deep understanding of modern model serving architectures and inference optimization techniques.
  • Strong expertise in writing GPU kernels for mobile devices.
  • Proficiency in designing robust evaluation frameworks for latency and memory constraints.
  • Knowledge of distributed inference systems, including Tensor, Pipeline, and Expert Parallelism.
  • Understanding of Diffusion Models, Vision Transformers, Pruning, Quantization, Flash attention, KV Cache, and Speculative Decoding.
  • Excellent English communication skills.

Responsibilities

  • Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage.
  • Build, run, and monitor controlled inference tests in both simulated and live production environments.
  • Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges.
  • Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring processing and memory metrics.
  • Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now