大模型榜單夾具變量風險

2026年8月26日 00:00
站內 AI 整理稿

Computer Science > Artificial Intelligence arXiv:2608.21382 (cs) [Submitted on 17 Jul 2026] Title:There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items Authors:V.S.

Raghu Parupudi View a PDF of the paper titled There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items, by V.S.

Raghu Parupudi View PDF HTML (experimental) Abstract:Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods.

Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next.

We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items.

We introduce the \textit{fragility grid}: 12 open-weight instruction-tuned LLMs from 4 families answer the same 3{,}679 items from 4 benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) under 26 equally defensible harness configurations, recording one correctness bit for every model, item, and configuration.

The comparison is matched, since the items, the weights, and the greedy decoding stay fixed while only the harness varies.Under the grid a model's score is a band rather than a point: gemma4-31b scores between 31 and 89 percent depending only on the harness.Three results follow.

On the items that two adjacent models both answer stably the pair is tied, and config-fragile items carry 95.7 percent of a pair's gap on average.Four of the 12 models reach rank one under some configuration, so the harness selects the winner.

Item discrimination, the property that benchmark-compression methods maximize, correlates with fragility at 0.28 (95 percent CI 0.25 to 0.30), so compression keeps the fragile items rather than removing them.

The scoring choice, not the option order that protocols usually fix, is the load-bearing axis.

We release the per-item records and the analysis script, from which every number regenerates on a CPU in seconds, and we position the fragility grid as a check a leaderboard can run before it reports an order.Comments: 25 pages in total Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2608.

21382 [cs.AI] (or arXiv:2608.21382v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.

21382 Focus to learn more arXiv-issued DOI via DataCite Submission history From: V Subrahmanya Raghu Ram Kishore Parupudi [view email] [v1] Fri, 17 Jul 2026 00:28:29 UTC (80 KB) Full-text links: Access Paper: View a PDF of the paper titled There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items, by V.

S.Raghu ParupudiView PDFHTML (experimental)TeX Source view license Current browse context: cs.AI < prev | next > new | recent | 2026-08 Change to browse by: cs References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...

Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.

ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?

) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.

AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

量子位生成式AI

Falcon TST 2.0獲世界權威測評第一名,推動時間序列基礎模型從通用預測走向金融應用

< img id="wx_img" src="https://www.qbitai.com/wp-content/uploads/imgs/qbitai-logo-1.png" width="400" height="400"> Falcon TST 2.0獲世界權威測評第一名,推動時間序列基礎模型從通用預測走向金融應用 量子位的朋友們 2026-08-26 11:25:00 來源:量子位 螞蟻國際日前正式發佈自研時序AI預測大模型“鷹序TST”2.0版。 螞蟻國際日前正式發佈自研時序AI預測大模型“鷹序TST”(FalconTST,Time-Series Transformer)2.0版。“鷹序TST”2.0在全球權威基準測試中取得最優成績(SOTA),平均絕對比例誤差(MASE)降至0.666,超越多家全球頂尖科技公司的時間序列預測模型。該模型專為跨境支付外匯風險管理打造,並將拓展至電商供應鏈需求預測、航空運營管理等更多行業場景。 目前,巴克萊銀行、花旗銀行、德意志銀行、渣打銀行等大型金融機構已將“鷹序TST”2.0應用於現金流預測與外匯管理,提升流動性風險敞口的預測能力。“鷹序TST”2.0的預測準確率已穩定超過93%,具備向物流、航空、電商等金融以外行業拓展的應用能力。 2025年以來,Amazon Chronos-2、Google TimesFM 2.5、Salesforce Moirai 2.0和IBM FlowState相繼把多變量、長上下文、概率預測、跨採樣率適配等能力推向前臺;與此同時,金融市場的低信噪比、機制切換、厚尾風險和跨市場聯動,使通用榜單優勢很難直接轉化為穩定的金融預測收益。螞蟻國際Falcon TST的演進具有代表性:2025年先在外匯敞口和資金管理中完成業務驗證,2026年又以Falcon-2.0強化單變量預測效率,以Falcon-X補足異構

剛剛
IT之家生成式AI

湯道生內部發文回應騰訊 AI 慢了:AI 競爭不是短跑,熬得久比起得早更重要

作者:汪淼 責編:汪淼 評論: 8 月 26 日消息,據鞭牛士今日報道,騰訊集團高級執行副總裁、雲與智慧產業事業群 CEO 湯道生 8 月 25 日晚在騰訊內部刊物《知點》年刊上發表一篇文章。湯道生表示,這篇文章由他和 AI 協同完成,AI 已經從單純的執行工具,變成輔助思考的協作夥伴。

剛剛
IT之家生成式AI

OpenAI 重啟 ChatGPT Plus 賬號五小時限額機制,Pro 賬號暫不受影響

作者:遠洋 責編:遠洋 評論: 感謝網友 肖戰姊姊、咩咩洋、我家小妞很拽、不一樣的體驗、這個暱稱不怎麼樣、秦淮一夢 的線索投遞!8 月 26 日消息,經過自 7 月 13 日開始為期六週的無每日使用額度限制後,OpenAI 正式恢復了針對標準版 ChatGPT Plus 用戶的嚴格“五小時使用上限”,該限制適用於使用 Codex 和 ChatGPT Work 環境的用戶。

剛剛