WEDNESDAY · 23 SEP 2026 · 21:19 UTCRSSMY BRIEFINGS
TOOLS23 SEPT 2026 · 09:34 UTCLIVE COVERAGE

UK AISI and EvalEval Publish Reproducible Benchmark Results

UK AI Security Institute and the EvalEval Coalition have published reproducible benchmark results, releasing Evaluation Cards for five core benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0-tested on six frontier LLMs and two cyber-focused evaluations, using the Every Eval Ever schema to ensure transparency.

EDITORIAL DESK
UK AISI and EvalEval Publish Reproducible Benchmark Results
UK AI Security Institute and the EvalEval Coalition have published reproducible benchmark results, releasing Evaluation