Machine Learning Engineer, Speech - Joint Audio-Video Modeling
New
C
Cantina LabsSocial AI
Remote (U.S. or Europe)Full-TimeSenior
SalaryThe anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000).
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PyTorchDeep LearningGenerative AI
Requirements
- Exceptional research and development experience with large-scale audio models (>8B parameters, >500k hours of data).
- Hands-on experience with diffusion or flow-matching transformers, including samplers, schedules, and conditioning.
- Hands-on experience training audio VAEs, neural audio codecs, and vocoders.
- Proficiency with multi-node, multi-GPU distributed training (FSDP, DeepSpeed, or equivalent).
- Strong software engineering skills with a track record of building complex, production-quality systems.
- Deep expertise in PyTorch and performance optimization (profiling, CUDA, Triton, or C++).
- Proven experience shipping large-scale speech, audio, or multimodal generative models to production.
- Background working with large-scale ML data and iterating on data quality.
- Experience with voice cloning, speech control, or expressive speech generation.
- Notable publications or open-source contributions in speech, audio, or ML.
Responsibilities
- Design, train, and improve audio VAEs, neural codecs, and vocoders for generative models.
- Architect, implement, and train diffusion and flow-matching transformers for large-scale audio and video generation.
- Design audio conditioning and cross-modal alignment inside joint audio-video models.
- Define data requirements and collaborate on acquisition, curation, and quality filtering for speech and AV corpora.
- Design objective and subjective evaluation metrics for fidelity, intelligibility, and AV-sync.
- Drive inference efficiency through distillation, quantization, and kernel optimization.
- Partner with infrastructure teams to manage distributed training and production deployment.
- Contribute to safety guardrails, watermarking, and misuse mitigation for generative models.
View Full Description & ApplyYou'll be redirected to the employer's site