Machine Learning Engineer - Model Evaluation & Experimentation

New
W
Weekday AIArtificial Intelligence
United StatesContractMiddle
Salary60 - 90 USD per hour
Apply NowOpens the employer's application page

Job Details

Experience
Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
Required Skills
PythonGitMachine LearningData scienceDeep LearningLLM

Requirements

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
  • Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Ability to commit approximately 35 hours per week on a consistent basis.
  • Familiarity with reinforcement learning concepts (preferred).
  • Experience with AI evaluation, benchmark development, AI training, or task authoring (highly desirable).

Responsibilities

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
  • Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes.
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior.
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
  • Collaborate with AI researchers and fellow subject matter experts to continuously improve benchmark quality, technical rigor, and evaluation consistency.
View Full Description & ApplyYou'll be redirected to the employer's site
60 - 90 USD per hour
Apply Now