Research Scientist, Video & Multimodal

I
InnodataData Engineering, AI
Remote - United StatesFull-TimeSenior
Salary160,000 - 185,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
Roughly 5+ years
Required Skills
Machine LearningPyTorchComputer Vision

Requirements

  • 5+ years of hands-on industry experience in video understanding or multimodal ML.
  • Bachelor's degree in computer science, electrical engineering, or related technical field.
  • Strong proficiency in PyTorch.
  • Experience with ffmpeg and decord pipelines.
  • Fluency in annotation formats (temporal, COCO-style) and data tools like WebDataset, Parquet, Arrow, and HuggingFace datasets.
  • Experience fine-tuning large video or vision-language models using HuggingFace transformers, PEFT, and efficient inference.
  • Experience with long-form video, streaming, temporal segmentation, or synthetic video generation.
  • Track record of first-author publications or open-source contributions at major venues (CVPR, ICCV, NeurIPS, etc.).
  • Ability to communicate complex modeling and data decisions to expert and non-expert audiences.

Responsibilities

  • Translate requirements for video understanding, temporal localization, and generation into data specifications.
  • Develop evaluation methodologies for temporal grounding, long-context reasoning, and cross-modal retrieval.
  • Design evaluation strategies for video generation models regarding fidelity and physical plausibility.
  • Structure and sample video data to optimize model value.
  • Execute experiments including fine-tuning and model evaluation to prove the impact of data decisions.
  • Design adversarial evaluations to identify system failures.
  • Publish research benchmarks and papers.
  • Collaborate with annotation teams and subject-matter experts to operationalize data collection.
View Full Description & ApplyYou'll be redirected to the employer's site
160,000 - 185,000 USD per year
Apply Now