Senior AI/ML Test and Evaluation Engineer
New
O
OpenTeamsArtificial Intelligence
Washington, DC; Denver, CO; or Colorado Springs, CO preferred (hybrid). Highly qualified candidates outside these locations may also be considered.Full-TimeSenior
Salary145,000 - 250,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years of software engineering or machine learning engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models
- Required Skills
- PythonMachine LearningPyTorchData science
Requirements
- 6+ years of software engineering or machine learning engineering experience
- 3+ years evaluating, benchmarking, or deploying ML models in production or applied research environments
- Strong Python proficiency in a machine learning or data science context
- Hands-on experience with common ML frameworks and tooling, such as PyTorch and the Hugging Face ecosystem
- Experience developing or using model evaluation harnesses, benchmark suites, or test and evaluation frameworks
- Experience designing evaluation metrics and applying appropriate statistical rigor when interpreting and reporting results
- Experience building repeatable and auditable evaluation pipelines with documented data provenance
- Experience evaluating large language models or agentic workflows using task-based, metric-based, or judgment-based scoring
- Strong written communication skills, including the ability to clearly explain evaluation methodologies, results, limitations, and failure modes
- Ability to work effectively in an evolving environment and translate mission needs into practical evaluation approaches
- Bachelor’s degree in computer science, mathematics, engineering, or a related field, or equivalent practical experience
Responsibilities
- Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows
- Develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment
- Curate and recommend candidate benchmarks based on mission needs and document the provenance of ground-truth and reference data
- Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including documented limitations and failure modes
- Define and contribute to common standards for benchmark expression, ingestion, and reporting
- Support partner organizations and vendors as they integrate their capabilities with shared evaluation standards
- Build lightweight expert-scoring workflows and measure inter-reviewer agreement for judgment-based evaluations
- Participate in structured feedback sessions with mission end users and incorporate findings into the platform and evaluation methodology
- Develop reference notebooks and example workflows that enable data-science-capable analysts to run, interpret, and extend evaluations
- Document technical approaches, evaluation results, and key decisions for Government stakeholders and internal teams
View Full Description & ApplyYou'll be redirected to the employer's site