AI Benchmark & Datasets Engineer / Researcher
New
P
PathwayArtificial Intelligence
Candidates based anywhere in the EU, UK, United States, and Canada will be considered.Full-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Required Skills
- Machine LearningData scienceGitHubLLM
Requirements
- Published at least one paper at NeurIPS, ICLR, or ICML as lead author or key contributor.
- Significantly contributed to a newsworthy LLM training effort.
- At least 6 months experience in a leading Machine Learning research center.
- ICPC World Finalist, or IOI, IMO, or IPhO medalist in High School.
- Experience with ML/LLM evaluation, data science, or technical product roles.
- Comfortable reading papers, leaderboards, and GitHub repositories.
- Ability to turn technical research into repeatable benchmark specifications.
- Experience translating technical detail into business value for cross-functional communication.
- Strong focus on high-quality data, reproducible experiments, and clear documentation.
- Fluent in English.
Responsibilities
- Proactively identify, prioritize, and curate relevant public and client-driven benchmarks across our target use cases and markets.
- Evaluate candidate benchmarks for clarity, data quality, evaluation methodology, and fit with our model roadmap.
- Run benchmarks with baseline models to validate setup, uncover edge cases, and de‑risk R&D runs.
- Hand off benchmark-ready packages to R&D including specs, data, evaluation scripts, expected metrics, and constraints.
- Maintain a shared vocabulary and documentation around benchmarks, datasets, and evaluation formats.
- Track and organize benchmark results, model leaderboards, and quality standards for different customers.
- Contribute to demos and public‑facing proof points based on benchmark outcomes.
View Full Description & ApplyYou'll be redirected to the employer's site