Bringing a Common Language to AI Evaluation
IBM, Thursday, July 23rd, 2026
IBM launches EveryEvalEver, a standardized format to make AI benchmark results consistent and reusable.
IBM and collaborators launched EveryEvalEver, a community-driven project establishing a standardized JSON format for reporting AI benchmark results.
It targets the widespread inconsistency in model evaluation, where identical tests can produce scores varying by up to 20 percentage points depending on the harness and reporting format.
The project's database already holds over 22,000 model results across 2,200 benchmarks. The goal is to lower evaluation costs, improve transparency, and make existing evaluation data easier to reuse.