Machine Learning Engineer, Speech - Joint Audio-Video Modeling

New
C
Cantina LabsSocial AI
Remote (U.S. or Europe)Full-TimeSenior
SalaryThe anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000).
Apply NowOpens the employer's application page

Job Details

Required Skills
PyTorchDeep LearningGenerative AI

Requirements

  • Exceptional research and development experience with large-scale audio models (>8B parameters, >500k hours of data).
  • Hands-on experience with diffusion or flow-matching transformers, including samplers, schedules, and conditioning.
  • Hands-on experience training audio VAEs, neural audio codecs, and vocoders.
  • Proficiency with multi-node, multi-GPU distributed training (FSDP, DeepSpeed, or equivalent).
  • Strong software engineering skills with a track record of building complex, production-quality systems.
  • Deep expertise in PyTorch and performance optimization (profiling, CUDA, Triton, or C++).
  • Proven experience shipping large-scale speech, audio, or multimodal generative models to production.
  • Background working with large-scale ML data and iterating on data quality.
  • Experience with voice cloning, speech control, or expressive speech generation.
  • Notable publications or open-source contributions in speech, audio, or ML.

Responsibilities

  • Design, train, and improve audio VAEs, neural codecs, and vocoders for generative models.
  • Architect, implement, and train diffusion and flow-matching transformers for large-scale audio and video generation.
  • Design audio conditioning and cross-modal alignment inside joint audio-video models.
  • Define data requirements and collaborate on acquisition, curation, and quality filtering for speech and AV corpora.
  • Design objective and subjective evaluation metrics for fidelity, intelligibility, and AV-sync.
  • Drive inference efficiency through distillation, quantization, and kernel optimization.
  • Partner with infrastructure teams to manage distributed training and production deployment.
  • Contribute to safety guardrails, watermarking, and misuse mitigation for generative models.
View Full Description & ApplyYou'll be redirected to the employer's site
The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000).
Apply Now