TOOLS23 SEPT 2026 · 09:34 UTC● LIVE COVERAGE
UK AISI and EvalEval Publish Reproducible Benchmark Results
UK AI Security Institute and the EvalEval Coalition have published reproducible benchmark results, releasing Evaluation Cards for five core benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0-tested on six frontier LLMs and two cyber-focused evaluations, using the Every Eval Ever schema to ensure transparency.
EDITORIAL DESK
09:29MODELSNew AI Models Promise Lower Costs for DevelopersEDITORIAL DESK09:23INDUSTRYMicrosoft takedown of AI-driven scam service that hit 12,000 accountsEDITORIAL DESK09:21TOOLSRabbit releases cloud-based AI agent that works on any PCEDITORIAL DESK09:20TOOLSUK AISI and EvalEval Share Reproducible AI Benchmark ResultsEDITORIAL DESK09:18MODELSParallel halves research time and cost with GPT-6 AstraEDITORIAL DESK



