Senior Data Scientist, AI Evaluation

New
J
JobgetherFinancial technology
CanadaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
Approximately 6–10 years of experience in quantitative data science, machine learning, or a related technical discipline, with focused experience in measurement, evaluation, experimentation, or model validation.
Required Skills
PythonSQL

Requirements

  • Approximately 6–10 years of experience in quantitative data science, machine learning, or a related technical discipline.
  • Focused experience in measurement, evaluation, experimentation, or model validation.
  • Strong quantitative and statistical foundations, including experience designing rigorous experiments and interpreting uncertainty measures.
  • Experience defining evaluation metrics, establishing ground truth, and assessing complex or ambiguous model outputs.
  • Proficiency in Python and SQL for data analysis, model assessment, and evaluation workflows.
  • Experience evaluating machine learning models in production and translating findings into practical improvements.
  • Ability to validate automated graders against human judgments and account for variability, non-determinism, and measurement bias.
  • A quantitative degree in data science, statistics, mathematics, computer science, or a related field is advantageous; equivalent practical experience is welcome.
  • Hands-on production evaluation experience with large language models or AI agents, evaluation harnesses, LLM-as-judge calibration, and continuous integration regression gates is a plus.
  • Experience evaluating text-to-SQL systems, analytics agents, or other AI applications with outputs verifiable against underlying data is advantageous.
  • Fintech, brokerage, financial services, or other high-consequence domain experience is desirable.

Responsibilities

  • Define ground truth, quality metrics, scoring methodologies, and evaluation frameworks for AI models and agents.
  • Build repeatable evaluation loops to monitor model quality, identify performance regressions, and support release decisions.
  • Apply statistical methods, including sample sizing, confidence intervals, significance testing, and methods for non-deterministic outputs.
  • Compare automated evaluation methods with human assessments and improve grader reliability and scoring accuracy.
  • Partner with Engineering and Analytics Engineering to operationalize evaluation frameworks and integrate testing into development workflows.
  • Analyze evaluation results, identify quality gaps, and recommend improvements to model and agent performance.
  • Establish evaluation guidelines, documentation practices, review processes, and quality standards.
  • Collaborate with Product, Engineering, Analytics Engineering, and business stakeholders to define success criteria and evaluation priorities.
  • Support informed AI release decisions through independent assessments of model performance and readiness.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now