AI Benchmark & Datasets Engineer / Researcher

New
P
PathwayArtificial Intelligence
Candidates based anywhere in the EU, UK, United States, and Canada will be considered.Full-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English
Required Skills
Machine LearningData scienceGitHubLLM

Requirements

  • Published at least one paper at NeurIPS, ICLR, or ICML as lead author or key contributor.
  • Significantly contributed to a newsworthy LLM training effort.
  • At least 6 months experience in a leading Machine Learning research center.
  • ICPC World Finalist, or IOI, IMO, or IPhO medalist in High School.
  • Experience with ML/LLM evaluation, data science, or technical product roles.
  • Comfortable reading papers, leaderboards, and GitHub repositories.
  • Ability to turn technical research into repeatable benchmark specifications.
  • Experience translating technical detail into business value for cross-functional communication.
  • Strong focus on high-quality data, reproducible experiments, and clear documentation.
  • Fluent in English.

Responsibilities

  • Proactively identify, prioritize, and curate relevant public and client-driven benchmarks across our target use cases and markets.
  • Evaluate candidate benchmarks for clarity, data quality, evaluation methodology, and fit with our model roadmap.
  • Run benchmarks with baseline models to validate setup, uncover edge cases, and de‑risk R&D runs.
  • Hand off benchmark-ready packages to R&D including specs, data, evaluation scripts, expected metrics, and constraints.
  • Maintain a shared vocabulary and documentation around benchmarks, datasets, and evaluation formats.
  • Track and organize benchmark results, model leaderboards, and quality standards for different customers.
  • Contribute to demos and public‑facing proof points based on benchmark outcomes.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now