AI QA Trainer - LLM Evaluation

New
M
Location: World Wide - RemoteContractSenior
Salary6 - 65 USD per hour
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonSQLRegression testing

Requirements

  • A bachelor’s, master’s, or PhD in computer science, data science, computational linguistics, statistics, or a related field is ideal.
  • Experience shipping QA for ML/AI systems is relevant.
  • Safety or red-team experience is relevant.
  • Experience with test automation frameworks such as PyTest is relevant.
  • Hands-on experience with LLM evaluation tooling such as OpenAI Evals, RAG evaluators, or W&B is relevant.
  • Skills in evaluation rubric design, adversarial testing/red-teaming, regression testing at scale, bias/fairness auditing, and grounding verification are valued.
  • Prompt and system-prompt engineering skills are valued.
  • Test automation experience with Python/SQL is valued.
  • Clear, metacognitive communication and high-signal bug reporting are essential.

Responsibilities

  • Evaluate model responses to real-world scenarios and evaluation prompts for factual accuracy and logical soundness.
  • Test for hallucinations, factual consistency, prompt-injection and jailbreak resistance, bias, reasoning reliability, tool-use correctness, and retrieval-augmentation fidelity.
  • Design and run test plans and regression suites.
  • Build evaluation rubrics and pass/fail criteria.
  • Capture reproducible error traces and document root-cause hypotheses.
  • Suggest improvements to prompt engineering, guardrails, and evaluation metrics.
  • Participate in adversarial red-teaming and automation using Python/SQL.
  • Create dashboards to track quality changes over time.
View Full Description & ApplyYou'll be redirected to the employer's site
6 - 65 USD per hour
Apply Now