Research Scientist, Video & Multimodal
I
InnodataData Engineering, AI
Remote - United StatesFull-TimeSenior
Salary160,000 - 185,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- Roughly 5+ years
- Required Skills
- Machine LearningPyTorchComputer Vision
Requirements
- 5+ years of hands-on industry experience in video understanding or multimodal ML.
- Bachelor's degree in computer science, electrical engineering, or related technical field.
- Strong proficiency in PyTorch.
- Experience with ffmpeg and decord pipelines.
- Fluency in annotation formats (temporal, COCO-style) and data tools like WebDataset, Parquet, Arrow, and HuggingFace datasets.
- Experience fine-tuning large video or vision-language models using HuggingFace transformers, PEFT, and efficient inference.
- Experience with long-form video, streaming, temporal segmentation, or synthetic video generation.
- Track record of first-author publications or open-source contributions at major venues (CVPR, ICCV, NeurIPS, etc.).
- Ability to communicate complex modeling and data decisions to expert and non-expert audiences.
Responsibilities
- Translate requirements for video understanding, temporal localization, and generation into data specifications.
- Develop evaluation methodologies for temporal grounding, long-context reasoning, and cross-modal retrieval.
- Design evaluation strategies for video generation models regarding fidelity and physical plausibility.
- Structure and sample video data to optimize model value.
- Execute experiments including fine-tuning and model evaluation to prove the impact of data decisions.
- Design adversarial evaluations to identify system failures.
- Publish research benchmarks and papers.
- Collaborate with annotation teams and subject-matter experts to operationalize data collection.
View Full Description & ApplyYou'll be redirected to the employer's site