AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: UK AISI And EvalEval Take On The Challenge Of Reproducible AI Benchmarks on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is publishing selected benchmark and cyber-evaluation results through EvalEval’s Evaluation Cards, which include verification, context and configuration information. The release accompanies an AISI paper on how inference-time compute and evaluation protocols affect results; it does not cover every AISI evaluation or provide a complete archive of underlying transcripts.

The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, which pair results with verification, evaluation context and configuration information, as detailed in the original analysis. The release accompanies an AISI paper examining how inference-time compute and evaluation protocols affect measured performance, giving researchers and other readers more detail for interpreting the results.

The paper’s main experiment reports results from five benchmarks across six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The benchmarks are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The release also includes AISI results from two cyber evaluations, Cyber CTFs and The Last Ones.

The cyber evaluations use a different, partly overlapping model set. The announcement does not enumerate that set, so the six models in the main experiment should not be assumed to cover the cyber results. EvalEval says the cards organize benchmark metadata, evaluation-run data and model metadata in a common format, and include verified results and setup information.

AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, examines the relationship between inference-time compute, evaluation protocols and scores. In its Humanity’s Last Exam analysis, the reported measure tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. The paper reports that, in runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased.

At a glance
announcementWhen: Announced alongside AISI’s paper, How I…
The developmentAISI has made selected evaluation results available through EvalEval’s Evaluation Cards alongside information about how the evaluations were run.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Evaluation Setup Matters

Benchmark scores are often cited as evidence of model capability, but a score can be difficult to interpret without knowing how the evaluation was run. AISI’s analysis of Humanity’s Last Exam illustrates that outcomes can vary with inference compute and feedback conditions. A result produced with repeated attempts and correctness feedback may represent different conditions from one produced without them.

Making setup information available alongside a result can help researchers and practitioners inspect individual runs and judge whether comparisons across models or studies are meaningful. It may also help policy teams and others who use evaluations as evidence about advanced AI capabilities. The cards provide a way to examine reported conditions; they do not, by themselves, determine which benchmark or protocol is most suitable.

The release addresses a reporting problem: evaluation results appear in different formats, and some reports may omit details needed to understand what a score represents. A shared record format can make those details easier to find. Its practical value will depend on what contributors document and how consistently they do so.

From AISI Workshop to Shared Records

The release builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current publication applies that infrastructure to selected AISI methods and findings that have been publicly reported where appropriate.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas including transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring evaluation results together with benchmark and model information. Those efforts address how results can be documented and inspected when repeating costly evaluations may not be feasible.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

What the Records Do Not Show

The announcement does not state how many records or transcripts are available, which setup fields appear for every benchmark, or whether independent researchers have reproduced the results. It says publicly reported methods and findings are being made available where appropriate, so the release should not be treated as a complete archive of AISI evaluations or underlying transcripts.

The announcement also does not list the models used in the two cyber evaluations, provide a date for each record, or describe how disagreements between results from different protocols would be handled. Those details would help readers assess coverage and compare reported runs. The existence of verification information on the cards does not establish that outside researchers have independently reproduced every result.

Broader Use of the EEE Schema

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Under its proposed workflow, model developers can submit verified results, while evaluation developers can report benchmark and run data using the EEE schema. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and review reporting practices across the collection.

No further release date or adoption milestone was specified. Broader use could make comparisons easier to inspect, but that will depend on the consistency and completeness of records contributors publish. The next visible developments are likely to be additional records and wider participation; the announcement does not set a timetable for either.

Key Questions

What has AISI released?

AISI is publishing selected benchmark and cyber-evaluation results through EvalEval’s Evaluation Cards, with verification, evaluation context and configuration information. The announcement does not describe a complete archive of its evaluations.

Which benchmarks are included in the main experiment?

The five are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The reported main experiment covers six models, while the cyber evaluations use a different, partly overlapping model set.

Why does inference-time compute matter?

AISI’s paper examines how additional inference compute and evaluation conditions relate to measured scores. Its Humanity’s Last Exam analysis reports that models given correctness feedback after attempts went on to solve additional tasks as token use increased.

Does the release include every AISI evaluation?

No such claim was made. The announcement describes selected publicly reported methods and findings being shared where appropriate and does not say that every evaluation or underlying transcript is included.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Choose a Computer for RStudio and Jupyter Work

Curious about selecting the perfect computer for RStudio and Jupyter? Discover essential features that can elevate your data analysis experience.

JMP Interface Fast‑Track Tutorial

Start exploring the JMP interface quickly with this fast-track tutorial that reveals essential tips to boost your data analysis skills.

Why Portable Monitors Are Better Than Most Students Expect

A surprising boost to student productivity, portable monitors offer comfort and versatility—discover how they can transform your study experience.

The Shift To SDL3: How Minecraft Java Edition Enhances Gaming Signals

Minecraft Java Edition now uses SDL3, marking a technical shift that impacts gaming signal monitoring and operator decision-making in real-time updates.