大模型榜單夾具變量風險
Computer Science > Artificial Intelligence arXiv:2608.21382 (cs) [Submitted on 17 Jul 2026] Title:There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items Authors:V.S.
Raghu Parupudi View a PDF of the paper titled There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items, by V.S.
Raghu Parupudi View PDF HTML (experimental) Abstract:Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods.
Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next.
We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items.
We introduce the \textit{fragility grid}: 12 open-weight instruction-tuned LLMs from 4 families answer the same 3{,}679 items from 4 benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) under 26 equally defensible harness configurations, recording one correctness bit for every model, item, and configuration.
The comparison is matched, since the items, the weights, and the greedy decoding stay fixed while only the harness varies.Under the grid a model's score is a band rather than a point: gemma4-31b scores between 31 and 89 percent depending only on the harness.Three results follow.
On the items that two adjacent models both answer stably the pair is tied, and config-fragile items carry 95.7 percent of a pair's gap on average.Four of the 12 models reach rank one under some configuration, so the harness selects the winner.
Item discrimination, the property that benchmark-compression methods maximize, correlates with fragility at 0.28 (95 percent CI 0.25 to 0.30), so compression keeps the fragile items rather than removing them.
The scoring choice, not the option order that protocols usually fix, is the load-bearing axis.
We release the per-item records and the analysis script, from which every number regenerates on a CPU in seconds, and we position the fragility grid as a check a leaderboard can run before it reports an order.Comments: 25 pages in total Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2608.
21382 [cs.AI] (or arXiv:2608.21382v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.
21382 Focus to learn more arXiv-issued DOI via DataCite Submission history From: V Subrahmanya Raghu Ram Kishore Parupudi [view email] [v1] Fri, 17 Jul 2026 00:28:29 UTC (80 KB) Full-text links: Access Paper: View a PDF of the paper titled There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items, by V.
S.Raghu ParupudiView PDFHTML (experimental)TeX Source view license Current browse context: cs.AI < prev | next > new | recent | 2026-08 Change to browse by: cs References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

Falcon TST 2.0獲世界權威測評第一名,推動時間序列基礎模型從通用預測走向金融應用
< img id="wx_img" src="https://www.qbitai.com/wp-content/uploads/imgs/qbitai-logo-1.png" width="400" height="400"> Falcon TST 2.0獲世界權威測評第一名,推動時間序列基礎模型從通用預測走向金融應用 量子位的朋友們 2026-08-26 11:25:00 來源:量子位 螞蟻國際日前正式發佈自研時序AI預測大模型“鷹序TST”2.0版。 螞蟻國際日前正式發佈自研時序AI預測大模型“鷹序TST”(FalconTST,Time-Series Transformer)2.0版。“鷹序TST”2.0在全球權威基準測試中取得最優成績(SOTA),平均絕對比例誤差(MASE)降至0.666,超越多家全球頂尖科技公司的時間序列預測模型。該模型專為跨境支付外匯風險管理打造,並將拓展至電商供應鏈需求預測、航空運營管理等更多行業場景。 目前,巴克萊銀行、花旗銀行、德意志銀行、渣打銀行等大型金融機構已將“鷹序TST”2.0應用於現金流預測與外匯管理,提升流動性風險敞口的預測能力。“鷹序TST”2.0的預測準確率已穩定超過93%,具備向物流、航空、電商等金融以外行業拓展的應用能力。 2025年以來,Amazon Chronos-2、Google TimesFM 2.5、Salesforce Moirai 2.0和IBM FlowState相繼把多變量、長上下文、概率預測、跨採樣率適配等能力推向前臺;與此同時,金融市場的低信噪比、機制切換、厚尾風險和跨市場聯動,使通用榜單優勢很難直接轉化為穩定的金融預測收益。螞蟻國際Falcon TST的演進具有代表性:2025年先在外匯敞口和資金管理中完成業務驗證,2026年又以Falcon-2.0強化單變量預測效率,以Falcon-X補足異構

深度實測「豆包工作」+飛書:目前最接近企業Agent終局的答案
豆包工作推出獨立電腦版並與飛書深度整合,實測能自動完成製圖、影片、網頁及採購比選等任務。其最大亮點是能直接讀取飛書中的群聊、文件與表格,理解企業上下文並整理選題、追蹤專案進度,被認為是目前最接近企業級Agent終局的答案。

湯道生內部發文回應騰訊 AI 慢了:AI 競爭不是短跑,熬得久比起得早更重要
作者:汪淼 責編:汪淼 評論: 8 月 26 日消息,據鞭牛士今日報道,騰訊集團高級執行副總裁、雲與智慧產業事業群 CEO 湯道生 8 月 25 日晚在騰訊內部刊物《知點》年刊上發表一篇文章。湯道生表示,這篇文章由他和 AI 協同完成,AI 已經從單純的執行工具,變成輔助思考的協作夥伴。

斯坦福研究:22~25 歲職場新人受 AI 衝擊最大,高等教育可緩衝影響
作者:故淵 責編:故淵 評論: 8 月 26 日消息,斯坦福數字經濟實驗室於 8 月 12 日更新論文,依託薪資服務商 ADP 覆蓋數百萬美國勞動者的高頻薪資數據,揭示生成式 AI 普及後勞動力市場的真實變化,報告提示,應屆與初入職場的年輕人正承受 AI 帶來的最直接衝擊。

Windows 任務管理器之父打造:Win11 版 TMOG 發佈,107 頁文檔讓 AI 編碼
Linux 系統發佈原生任務管理器 TMOG,最新版本號為 0.1.1。普拉默在推文中表示,1995 年在開發 Windows 任務管理器時,使用一臺 P90 配置的電腦,就是在 CPU 使用率不超過 2% 的情況下,實現更新、滾動和繪圖功能。

OpenAI 重啟 ChatGPT Plus 賬號五小時限額機制,Pro 賬號暫不受影響
作者:遠洋 責編:遠洋 評論: 感謝網友 肖戰姊姊、咩咩洋、我家小妞很拽、不一樣的體驗、這個暱稱不怎麼樣、秦淮一夢 的線索投遞!8 月 26 日消息,經過自 7 月 13 日開始為期六週的無每日使用額度限制後,OpenAI 正式恢復了針對標準版 ChatGPT Plus 用戶的嚴格“五小時使用上限”,該限制適用於使用 Codex 和 ChatGPT Work 環境的用戶。