Senior AI Systems Quality Engineer
New
J
JobgetherHealthcare AI
Based in United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years of software engineering experience, primarily focused on backend or platform systems.
- Required Skills
- AWSPythonMLFlowTypeScriptCI/CDDatabricks
Requirements
- Have 7+ years of software engineering experience, primarily focused on backend or platform systems.
- Have production experience designing and implementing automated AI testing and validation solutions.
- Have experience building custom testing, validation, or evaluation frameworks for complex and distributed systems.
- Have strong proficiency in Python and/or TypeScript within modern AI engineering environments.
- Have hands-on experience with LLM-based or agentic systems and non-deterministic behavior.
- Have experience designing AI testing at scale, including regression frameworks, long-tail evaluations, and broad test coverage.
- Understand CI/CD practices and have experience embedding automated tests and quality gates into deployment pipelines.
- Have solid knowledge of AWS cloud-native architectures.
- Have experience engineering for quality, reliability, governance, safety, and operational resilience.
- Have working knowledge of security, privacy, and operational risk in regulated or mission-critical environments.
- Have experience with AI testing methods such as non-deterministic output evaluation, drift detection, bias and fairness testing, and regression strategies.
- Be able to establish trust thresholds and operationalize metrics such as query accuracy, hallucination limits, explainability, and PHI-safe behavior as release criteria.
Responsibilities
- Build and deploy automated validation frameworks, test harnesses, and evaluation pipelines across the AI development lifecycle.
- Develop an AI testing platform integrated with Databricks and MLflow for repeatable testing, traceability, lineage, and auditability.
- Create large-scale, scenario-based test suites covering edge cases, long-tail scenarios, and system failure modes.
- Validate agentic orchestration behaviors, including tool usage, memory, decision logic, and non-deterministic outputs.
- Define system contracts, guardrails, safe-degradation patterns, and validation requirements at key system boundaries.
- Define measurable LLM quality signals, including grounding, hallucination rates, relevance, latency, cost, accuracy, and explainability.
- Integrate automated quality gates into CI/CD pipelines and run validation after model, prompt, or code changes.
- Build reusable testing libraries, frameworks, and components for consistent AI quality practices.
- Establish release-readiness criteria and support go/no-go decisions based on quality thresholds.
- Partner with AI, platform, security, and delivery teams to evaluate reliability, security, privacy, and operational risk.
View Full Description & ApplyYou'll be redirected to the employer's site