AI QA Trainer - LLM Evaluation
New
M
MeridialAI/ML
Location: World Wide - RemoteContractSenior
Salary6 - 65 USD per hour
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonSQLRegression testing
Requirements
- A bachelor’s, master’s, or PhD in computer science, data science, computational linguistics, statistics, or a related field is ideal.
- Experience shipping QA for ML/AI systems is relevant.
- Safety or red-team experience is relevant.
- Experience with test automation frameworks such as PyTest is relevant.
- Hands-on experience with LLM evaluation tooling such as OpenAI Evals, RAG evaluators, or W&B is relevant.
- Skills in evaluation rubric design, adversarial testing/red-teaming, regression testing at scale, bias/fairness auditing, and grounding verification are valued.
- Prompt and system-prompt engineering skills are valued.
- Test automation experience with Python/SQL is valued.
- Clear, metacognitive communication and high-signal bug reporting are essential.
Responsibilities
- Evaluate model responses to real-world scenarios and evaluation prompts for factual accuracy and logical soundness.
- Test for hallucinations, factual consistency, prompt-injection and jailbreak resistance, bias, reasoning reliability, tool-use correctness, and retrieval-augmentation fidelity.
- Design and run test plans and regression suites.
- Build evaluation rubrics and pass/fail criteria.
- Capture reproducible error traces and document root-cause hypotheses.
- Suggest improvements to prompt engineering, guardrails, and evaluation metrics.
- Participate in adversarial red-teaming and automation using Python/SQL.
- Create dashboards to track quality changes over time.
View Full Description & ApplyYou'll be redirected to the employer's site