Skip to Main Content

Breadcrumb

Introduction

The Artificial Intelligence Evaluation Center (AIEC) has been established to promote localized AI evaluation and third-party certification in Taiwan, thereby strengthening the development of trusted AI within the industry. The Center will periodically publish benchmark evaluation results for language models. In addition to adopting indicators based on the Chinese Language and Social Studies sections of the national high school entrance examination, AIEC also incorporates evaluation criteria reflecting Taiwanese values, aligning with global trends in AI sovereignty. These benchmarks serve as key references for developing locally adapted models or fine-tuning international models.

✪Sorted by the region of the developing organization: light orange represents European models, light blue represents U.S. models, light green represents local models, light purple represents Chinese models, and light beige represents Asian models.

✪Explanation of percentage values: figures above 50% are marked in black; figures below 50% are marked in red.

✪This round of testing results includes the addition of 2 small-scale models and 12 large-scale models.

✪Starting from May 2026, the testing results will include an additional field indicating the model testing date for reference.

  • Language Model Benchmark / Small Models (13B and below)

 
 

  • Language Model Benchmark / Large Models (above 13B)

 
 

Downloads:
Test Results of the July 2026 OpenSource Models(Large and Small Models).ods
Test Results of the July 2026 OpenSource Models(Large and Small Models).xlsx