On July 27, 2026, the US National Institute of Standards and Technology (NIST) announced the "AI Technology Evaluation (AITE)" program. The AITE program establishes a standardized and isolated testbed to conduct rigorous, trustworthy, and independent evaluations of model performance across diverse datasets, modalities, and specialized domains.
The AITE program employs a blind-testing mechanism within isolated environments. Models are evaluated on data they have never encountered before. This ensures that the evaluation results reflect a model's true generalization capabilities and performance, rather than its ability to memorize training data. NIST will provide data, metrics and scoring to help developers understand the performance of their models. NIST will regularly publish leaderboards and reports.
The AITE program adopts a dual-track parallel mechanism to incentivize participation from industry, academia, and research institutions. For Data Providers, organizations can submit unreleased, domain-specific datasets and targeted tasks. They receive empirical data detailing how top-tier models perform in their specific domain. For Model Providers, developers can submit their models for testing while safeguarding intellectual property and preventing data leakage. They receive precise benchmarking against industry peers. This two-way interaction fosters continuous improvement between high-quality data and advanced models.
In its early phase, the initiative focuses on the performance of large Vision-Language Models (VLMs) in specialized image analysis. The first wave covers three key domains: quantum science, genomics, and public safety. The program will gradually expand its task types, modalities, domains, and evaluation techniques. It will also scale its evaluation capabilities by increasing the number of test tasks and models. The initiative aims to eventually cover hundreds of tasks and domains. It will integrate security protections for Controlled Unclassified Information (CUI), achieve full automation in evaluations, and adopt innovative metrological methods. Ultimately, the program seeks to establish a long-term, dynamic evaluation ecosystem that continuously adapts to emerging technological needs.