英國AISI與EvalEval如何讓基準測試結果可重現
Back to Articles How UK AISI and EvalEval Are Making Benchmark Results Reproducible Published September 22, 2026 Update on GitHub Upvote 1 Avijit Ghosh evijit Follow evaleval Jenny Chim j-chim Follow evaleval Deep Joshi deeplumiere Follow evaleval Srishti srishtiy Follow evaleval Matt Kennedy wmmkennedy Follow evaleval Irene Solaiman irenesolaiman Follow evaleval Jessica McFadyen mcfadyen-aisi Follow ai-safety-institute Lynn Tan lynn-aisi Follow ai-safety-institute Coz coz-aisi Follow ai-safety-institute The EvalEval Coalition is thrilled to share that the UK AI Security Institute (AISI) is using EvalEval's infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.
AISI and EvalEval have previously collaborated on research that began at a joint workshop alongside NeurIPS 2025, and feedback from the Institute has helped shape the Every Eval Ever (EEE) schema.This next phase of the collaboration puts that shared infrastructure into practice.
Why reproducible evaluation reporting matters As AI deployment accelerates, evaluations are becoming increasingly important sources of evidence about model and system performance.Yet results are reported across many formats, platforms, and outlets, often without enough information to reproduce them.
Running the evaluations again may itself be prohibitively expensive.
EvalEval's mission is to improve this ecosystem through a shared reporting schema, Every Eval Ever, and an open platform, Evaluation Cards, that brings evaluation results and the information needed to interpret them into a common structure.
This builds naturally on AISI's work to make evaluation more efficient through OptStop, more statistically rigorous through HiBayES, and more standardised in areas including transcript analysis and capability elicitation.
Together, AISI and EvalEval are working to diagnose gaps in evaluation reporting and build shared infrastructure to close them.What AISI is sharing Transcript-level transparency matters not only for reproducibility, but also for analysis and diagnosis.
In this new phase of the collaboration, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate.
The release includes verified results, context, and configuration information for the five benchmarks in the paper's main experiment: HealthBench FrontierMath Humanity's Last Exam SWE-Bench Pro Terminal-Bench 2.0 These results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.
6, GPT-5, GPT-5.2, and GPT-5.4.The release also includes results from two related cyber evaluations—Cyber CTFs and The Last Ones—which use a different, partially overlapping set of models.
The data accompany AISI's paper, How Inference Compute Shapes Frontier LLM Evaluation, which studies how benchmark performance depends on inference-time compute and evaluation protocol.Performance on Humanity's Last Exam changes with evaluation protocol and inference compute.
Each curve shows the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task.When models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased.
When results are openly released with setup information, researchers and practitioners can examine individual studies more closely and compare findings across the wider ecosystem.
Where other reports lack these details, releases like AISI's provide verified reference points for interpreting evaluations in context—for example, by helping researchers understand how setup choices may influence reported performance.
As more evaluators adopt EEE, open comparisons like these can support broader and more reliable meta-research.AISI's Terminal-Bench 2.0 results alongside other reported evaluations for the same models, under different evaluation setups.
We are excited about this adoption and look forward to further standardising and sharing evaluations with AISI and other AI evaluation organisations.Contribute to the shared mission Model developers: Report verified evaluation results.
Evaluation developers: Report benchmarks and run data using the Every Eval Ever schema.Evaluation, governance, and policy researchers: Explore Evaluation Cards by benchmark or model, or use it to examine the state of evaluation reporting as a whole.
About the EvalEval Coalition The EvalEval Coalition is a research community developing scientifically grounded research and robust deployment infrastructure for the evaluation ecosystem.
Its goal is to improve evaluation science, address the lack of consensus around documenting evaluation applicability and utility, and broaden coverage of the impacts that matter for scientific research and policy analysis.
The coalition's flagship projects include Every Eval Ever, a shared schema and repository for evaluation results, and Evaluation Cards, which combines benchmark metadata, evaluation-run data, and model metadata into interpretable records.
Together, they make it easier to understand when apparently similar scores were produced under meaningfully different conditions.About the UK AI Security Institute The UK AI Security Institute is a research organisation within the UK government's Department for Science, Innovation and Technology.
Its mission is to equip governments with a scientific understanding of the risks posed by advanced AI.AISI conducts research and builds infrastructure to understand advanced AI capabilities and impacts, develop and test mitigations, and inform policy.
Further reading How Inference Compute Shapes Frontier LLM Evaluation HiBayES: Improving LLM evaluation with hierarchical Bayesian modelling HiBayES paper OptStop paper Every Eval Ever Evaluation Cards More Articles from our Blog evaluationcommunityleaderboard Featuring Every Eval Ever Results on Hugging Face Model Pages +3 54 June 30, 2026 nlpevaluationretrieval Introducing RTEB: A New Standard for Retrieval Evaluation +2 149 October 1, 2025 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 1
Related
相關文章

聯合國 AI 專家反對 AI 末日論,呼籲理性討論
作者:遠洋 責編:遠洋 評論: 9 月 22 日消息,聯合國下屬人工智能委員會的多名專家於當地時間週一發出提醒,反對業內部分從業者針對 AI 風險所發表的世界末日式言論。獨立機構國際 AI 科學專家組(International Scientific Panel on AI)成員、谷歌 DeepMind 高管若埃爾 · 巴拉爾(Joelle Barral)表示:“身為科研工作者,我們應當投入精力開展科學研究,釐清哪些內容已有定論、哪些尚且屬於未知。
在線蒸餾顯示小樣本也能傳遞推理方式
Computer Science > Machine Learning arXiv:2609.14193 (cs) [Submitted on 12 Sep 2026 (v1), last revised 17 Sep 2026 (this version, v2)] Title:Data-free On-policy Distillation Authors:Gengsheng Li, Mao Zheng, Mingyang Song。

宇樹科技發佈 Dex5-S 靈巧手:22 自由度、真手 1:1 尺寸,售價 3.99 萬元起
USB×1,通信波特率 6Mbps;控制頻率方面,T1s 以太網 / USB 最高 1000Hz,高速 485 為 50Hz。感知反饋涵蓋關節模式、位置、速度、力矩、溫度、電壓電流等,控制指令則包括關節模式、位置、速度、力矩、剛度係數與阻尼係數。

華為汪濤:華為要打造AI算力底座,只做好一顆芯片遠遠不夠
。 昇騰960DT研發進度比原計劃提前三個季度,單芯片算力實現倍增,960DT支持2 PFLOPS FP8、4 PFLOPS FP4,HBM容量最高288GB、帶寬9.6TB/s。 對應的昇騰960超節點做到4096張NPU卡,最高8EFLOPS FP8算力、超過1PB HBM容量。

OpenAI 一次性攤開六份失準報告:模型越界不再是偶然,而是三種可復現的機制
這些報告合在一起,勾勒出的不是個別翻車,而是三類反覆出現的"越界機制"。第一類發生在任務交接的縫隙裡。模型在寫給下一步的摘要中,會擅自添加指令、或者隱瞞原本的要求,從而悄悄改變後續任務的走向——問題不在最終交付物,而在那段承上啟下的文字被人忽略。
OpenAI 一次性攤開六份失準報告:模型越界不再是偶然,而是三種可復現的機制
這些報告合在一起,勾勒出的不是個別翻車,而是三類反覆出現的"越界機制"。第一類發生在任務交接的縫隙裡。模型在寫給下一步的摘要中,會擅自添加指令、或者隱瞞原本的要求,從而悄悄改變後續任務的走向——問題不在最終交付物,而在那段承上啟下的文字被人忽略。