AI Safety Expert

New
J
JobgetherAI Safety
CanadaContract
Salary48 - 62 CAD per hour
Apply NowOpens the employer's application page

Job Details

Languages
English and Norwegian
Required Skills
CybersecurityPrompt Engineering

Requirements

  • Fluent or professional-level proficiency in both English and Norwegian, with strong written and verbal communication skills.
  • Prior experience in AI red teaming, adversarial AI work, cybersecurity, penetration testing, or socio-technical risk assessment.
  • Demonstrated ability to systematically probe complex systems for vulnerabilities, unexpected behaviors, and potential misuse.
  • Experience applying structured frameworks, taxonomies, benchmarks, or testing methodologies to security or safety evaluations.
  • Strong analytical and critical-thinking skills, combined with the creativity needed to develop novel probing strategies.
  • Ability to clearly document technical findings and communicate them effectively to both technical and non-technical audiences.
  • Strong attention to detail and commitment to producing reproducible, high-quality evaluation results.
  • Adaptability and ability to move efficiently between different projects, models, use cases, and customer requirements.
  • Experience in adversarial machine learning, cybersecurity, or socio-technical risk is preferred.
  • Creative-probing skills developed through psychology, acting, creative writing, or related disciplines are an advantage.

Responsibilities

  • Conduct red-team testing of conversational AI models and agents to identify jailbreaks, prompt-injection vulnerabilities, misuse scenarios, and opportunities for bias exploitation.
  • Develop creative and realistic adversarial prompts and attack scenarios designed to probe model weaknesses.
  • Generate high-quality human evaluation data by annotating model failures, classifying vulnerabilities, and identifying systemic safety risks.
  • Apply established taxonomies, benchmarks, testing frameworks, and playbooks to ensure consistent and rigorous evaluations.
  • Document findings in a reproducible manner through detailed reports, datasets, attack cases, and supporting evidence.
  • Clearly communicate technical findings and risk assessments to both technical and non-technical stakeholders.
  • Adapt testing approaches across different projects, AI systems, use cases, and customer requirements while maintaining consistent quality standards.
  • Contribute insights that help improve AI safety practices, model robustness, and risk mitigation strategies.
View Full Description & ApplyYou'll be redirected to the employer's site
48 - 62 CAD per hour
Apply Now