Adversarial AI Specialist
New
M
MercorAI Research
United StatesContractMiddle
Salary20 - 22 USD per hour
Apply NowOpens the employer's application page
Job Details
- Languages
- English, Urdu
- Required Skills
- Cybersecurity
Requirements
- Native fluency in English and Urdu.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Strong communication skills for explaining risks to technical and non-technical stakeholders.
- Experience in Adversarial ML (e.g., jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction) preferred.
- Background in Cybersecurity (e.g., penetration testing, exploit development, reverse engineering) preferred.
- Expertise in socio-technical risk, including harassment/disinfo probing and abuse analysis.
- Creative probing skills in psychology, acting, or writing for unconventional adversarial thinking.
Responsibilities
- Red team conversational AI models and agents by conducting jailbreaks, prompt injections, misuse cases, and bias exploitation.
- Generate high-quality human data, annotate model failures, classify vulnerabilities, and flag systemic risks.
- Apply structure using taxonomies, benchmarks, and playbooks to ensure testing consistency.
- Document findings reproducibly by producing reports, datasets, and attack cases.
- Work independently and asynchronously to meet deadlines and contribute to improved AI model performance.
View Full Description & ApplyYou'll be redirected to the employer's site