Senior AI/ML Test and Evaluation Engineer

New
O
OpenTeamsArtificial Intelligence
Washington, DC; Denver, CO; or Colorado Springs, CO preferred (hybrid). Highly qualified candidates outside these locations may also be considered.Full-TimeSenior
Salary145,000 - 250,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
6+ years of software engineering or machine learning engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models
Required Skills
PythonMachine LearningPyTorchData science

Requirements

  • 6+ years of software engineering or machine learning engineering experience
  • 3+ years evaluating, benchmarking, or deploying ML models in production or applied research environments
  • Strong Python proficiency in a machine learning or data science context
  • Hands-on experience with common ML frameworks and tooling, such as PyTorch and the Hugging Face ecosystem
  • Experience developing or using model evaluation harnesses, benchmark suites, or test and evaluation frameworks
  • Experience designing evaluation metrics and applying appropriate statistical rigor when interpreting and reporting results
  • Experience building repeatable and auditable evaluation pipelines with documented data provenance
  • Experience evaluating large language models or agentic workflows using task-based, metric-based, or judgment-based scoring
  • Strong written communication skills, including the ability to clearly explain evaluation methodologies, results, limitations, and failure modes
  • Ability to work effectively in an evolving environment and translate mission needs into practical evaluation approaches
  • Bachelor’s degree in computer science, mathematics, engineering, or a related field, or equivalent practical experience

Responsibilities

  • Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows
  • Develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment
  • Curate and recommend candidate benchmarks based on mission needs and document the provenance of ground-truth and reference data
  • Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including documented limitations and failure modes
  • Define and contribute to common standards for benchmark expression, ingestion, and reporting
  • Support partner organizations and vendors as they integrate their capabilities with shared evaluation standards
  • Build lightweight expert-scoring workflows and measure inter-reviewer agreement for judgment-based evaluations
  • Participate in structured feedback sessions with mission end users and incorporate findings into the platform and evaluation methodology
  • Develop reference notebooks and example workflows that enable data-science-capable analysts to run, interpret, and extend evaluations
  • Document technical approaches, evaluation results, and key decisions for Government stakeholders and internal teams
View Full Description & ApplyYou'll be redirected to the employer's site
145,000 - 250,000 USD per year
Apply Now