AI SDET (Software Development Engineer in Test)

New
J
JobgetherAI / Software Quality
BrazilFull-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
Python

Requirements

  • Advanced proficiency in Python 3.11+, Pytest, and modern automated software testing frameworks.
  • Hands-on experience with AI evaluation platforms and methodologies, including Langfuse, RAGAS, or comparable tools.
  • Practical knowledge of LLM-as-a-judge evaluation patterns and approaches for assessing non-deterministic AI outputs.
  • Understanding of model risk management principles and regulatory compliance testing for AI-driven systems.
  • Experience with shadow-mode deployment strategies, statistical validation, and comparative model evaluation.
  • Strong understanding of automated testing principles and the ability to design robust, repeatable test scenarios.
  • Exceptional attention to detail, particularly when producing evidence related to regulatory requirements and quality gates.
  • A naturally skeptical and adversarial mindset, with the ability to question expected behavior and proactively identify edge cases.
  • Strong analytical and problem-solving abilities, with a structured approach to investigating complex AI behavior.
  • Excellent collaboration and communication skills, particularly when working with engineers to diagnose and resolve logic failures.
  • Ability to work effectively in an environment where AI systems, testing methodologies, and requirements evolve rapidly.

Responsibilities

  • Own and manage the AI evaluation harness, including the creation, maintenance, and continuous improvement of automated evaluation workflows.
  • Develop and maintain “Golden Path” scenarios representing expected AI behavior and critical user journeys.
  • Design adversarial and edge-case test scenarios to identify weaknesses in AI agents and expose failures that conventional testing may overlook.
  • Validate regulatory quality gates and compliance requirements, including frameworks such as SR 26-2 and NYDFS Regulation 187.
  • Execute shadow-mode validation to compare AI system performance against established human benchmarks and analyze differences statistically.
  • Test for data quality degradation, population segmentation errors, inconsistencies, and other risks that could affect model reliability.
  • Evaluate the reliability of human-in-the-loop checkpoints and verify that required intervention points operate correctly.
  • Confirm that AI-generated outputs consistently meet defined standards, including explanation blocks, signal types, and other required output structures.
  • Produce clear, traceable evaluation evidence and communicate findings to engineering and other technical stakeholders.
  • Partner closely with engineers to investigate failures, identify root causes, and improve AI system logic and quality.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now