Machine Learning Engineer - Voice Conversion
New
C
CantinaAI Social Platform
Remote (U.S. or Europe)Full-TimeMiddle
Salary$200,000-$220,000 (€170,000-€190,000)
Apply NowOpens the employer's application page
Job Details
- Required Skills
- Machine LearningPyTorchC++Generative AI
Requirements
- Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).
- Deep hands-on experience with diffusion and/or flow-matching transformers.
- Practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
- Expertise in training audio VAEs, neural audio codecs, and vocoder latent/tokenizer design.
- Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed).
- Strong software engineering skills with a track record of building complex systems.
- Advanced proficiency with PyTorch, profiling, and performance optimization (CUDA/Triton/C++).
- Proven experience shipping large-scale speech/audio or multimodal generative models to production.
- Ability to iterate on large-scale ML data using subjective and objective quality signals.
- Experience with voice cloning, speech control, or expressive speech generation.
- Notable publications or open-source contributions in speech/audio/ML.
Responsibilities
- Architect, implement, pre-train, fine-tune, and align large-scale speech models.
- Design and execute scientific experiments to advance model performance.
- Develop internal tools to improve team productivity.
- Define data requirements and lead strategies for acquisition, curation, and synthetic data.
- Design rigorous objective and subjective evaluation frameworks, including listening tests and robustness checks.
- Harden the training-to-inference pipeline to meet production SLAs, including latency and cost profiling.
- Implement safety and consent guardrails for responsible speech technology.
View Full Description & ApplyYou'll be redirected to the employer's site