AI SDET (Software Development Engineer in Test)
New
J
JobgetherAI / Software Quality
BrazilFull-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- Python
Requirements
- Advanced proficiency in Python 3.11+, Pytest, and modern automated software testing frameworks.
- Hands-on experience with AI evaluation platforms and methodologies, including Langfuse, RAGAS, or comparable tools.
- Practical knowledge of LLM-as-a-judge evaluation patterns and approaches for assessing non-deterministic AI outputs.
- Understanding of model risk management principles and regulatory compliance testing for AI-driven systems.
- Experience with shadow-mode deployment strategies, statistical validation, and comparative model evaluation.
- Strong understanding of automated testing principles and the ability to design robust, repeatable test scenarios.
- Exceptional attention to detail, particularly when producing evidence related to regulatory requirements and quality gates.
- A naturally skeptical and adversarial mindset, with the ability to question expected behavior and proactively identify edge cases.
- Strong analytical and problem-solving abilities, with a structured approach to investigating complex AI behavior.
- Excellent collaboration and communication skills, particularly when working with engineers to diagnose and resolve logic failures.
- Ability to work effectively in an environment where AI systems, testing methodologies, and requirements evolve rapidly.
Responsibilities
- Own and manage the AI evaluation harness, including the creation, maintenance, and continuous improvement of automated evaluation workflows.
- Develop and maintain “Golden Path” scenarios representing expected AI behavior and critical user journeys.
- Design adversarial and edge-case test scenarios to identify weaknesses in AI agents and expose failures that conventional testing may overlook.
- Validate regulatory quality gates and compliance requirements, including frameworks such as SR 26-2 and NYDFS Regulation 187.
- Execute shadow-mode validation to compare AI system performance against established human benchmarks and analyze differences statistically.
- Test for data quality degradation, population segmentation errors, inconsistencies, and other risks that could affect model reliability.
- Evaluate the reliability of human-in-the-loop checkpoints and verify that required intervention points operate correctly.
- Confirm that AI-generated outputs consistently meet defined standards, including explanation blocks, signal types, and other required output structures.
- Produce clear, traceable evaluation evidence and communicate findings to engineering and other technical stakeholders.
- Partner closely with engineers to investigate failures, identify root causes, and improve AI system logic and quality.
View Full Description & ApplyYou'll be redirected to the employer's site