Senior AI Systems Quality Engineer

New
J
JobgetherHealthcare AI
Based in United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
7+ years of software engineering experience, primarily focused on backend or platform systems.
Required Skills
AWSPythonMLFlowTypeScriptCI/CDDatabricks

Requirements

  • Have 7+ years of software engineering experience, primarily focused on backend or platform systems.
  • Have production experience designing and implementing automated AI testing and validation solutions.
  • Have experience building custom testing, validation, or evaluation frameworks for complex and distributed systems.
  • Have strong proficiency in Python and/or TypeScript within modern AI engineering environments.
  • Have hands-on experience with LLM-based or agentic systems and non-deterministic behavior.
  • Have experience designing AI testing at scale, including regression frameworks, long-tail evaluations, and broad test coverage.
  • Understand CI/CD practices and have experience embedding automated tests and quality gates into deployment pipelines.
  • Have solid knowledge of AWS cloud-native architectures.
  • Have experience engineering for quality, reliability, governance, safety, and operational resilience.
  • Have working knowledge of security, privacy, and operational risk in regulated or mission-critical environments.
  • Have experience with AI testing methods such as non-deterministic output evaluation, drift detection, bias and fairness testing, and regression strategies.
  • Be able to establish trust thresholds and operationalize metrics such as query accuracy, hallucination limits, explainability, and PHI-safe behavior as release criteria.

Responsibilities

  • Build and deploy automated validation frameworks, test harnesses, and evaluation pipelines across the AI development lifecycle.
  • Develop an AI testing platform integrated with Databricks and MLflow for repeatable testing, traceability, lineage, and auditability.
  • Create large-scale, scenario-based test suites covering edge cases, long-tail scenarios, and system failure modes.
  • Validate agentic orchestration behaviors, including tool usage, memory, decision logic, and non-deterministic outputs.
  • Define system contracts, guardrails, safe-degradation patterns, and validation requirements at key system boundaries.
  • Define measurable LLM quality signals, including grounding, hallucination rates, relevance, latency, cost, accuracy, and explainability.
  • Integrate automated quality gates into CI/CD pipelines and run validation after model, prompt, or code changes.
  • Build reusable testing libraries, frameworks, and components for consistent AI quality practices.
  • Establish release-readiness criteria and support go/no-go decisions based on quality thresholds.
  • Partner with AI, platform, security, and delivery teams to evaluate reliability, security, privacy, and operational risk.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now