AI Safety Expert
New
J
JobgetherAI Safety
CanadaContract
Salary48 - 62 CAD per hour
Apply NowOpens the employer's application page
Job Details
- Languages
- English and Norwegian
- Required Skills
- CybersecurityPrompt Engineering
Requirements
- Fluent or professional-level proficiency in both English and Norwegian, with strong written and verbal communication skills.
- Prior experience in AI red teaming, adversarial AI work, cybersecurity, penetration testing, or socio-technical risk assessment.
- Demonstrated ability to systematically probe complex systems for vulnerabilities, unexpected behaviors, and potential misuse.
- Experience applying structured frameworks, taxonomies, benchmarks, or testing methodologies to security or safety evaluations.
- Strong analytical and critical-thinking skills, combined with the creativity needed to develop novel probing strategies.
- Ability to clearly document technical findings and communicate them effectively to both technical and non-technical audiences.
- Strong attention to detail and commitment to producing reproducible, high-quality evaluation results.
- Adaptability and ability to move efficiently between different projects, models, use cases, and customer requirements.
- Experience in adversarial machine learning, cybersecurity, or socio-technical risk is preferred.
- Creative-probing skills developed through psychology, acting, creative writing, or related disciplines are an advantage.
Responsibilities
- Conduct red-team testing of conversational AI models and agents to identify jailbreaks, prompt-injection vulnerabilities, misuse scenarios, and opportunities for bias exploitation.
- Develop creative and realistic adversarial prompts and attack scenarios designed to probe model weaknesses.
- Generate high-quality human evaluation data by annotating model failures, classifying vulnerabilities, and identifying systemic safety risks.
- Apply established taxonomies, benchmarks, testing frameworks, and playbooks to ensure consistent and rigorous evaluations.
- Document findings in a reproducible manner through detailed reports, datasets, attack cases, and supporting evidence.
- Clearly communicate technical findings and risk assessments to both technical and non-technical stakeholders.
- Adapt testing approaches across different projects, AI systems, use cases, and customer requirements while maintaining consistent quality standards.
- Contribute insights that help improve AI safety practices, model robustness, and risk mitigation strategies.
View Full Description & ApplyYou'll be redirected to the employer's site