BottleCap AI 發布 ThinkingCap-Qwen3.8-27B:思考 token 減少 37.2%,準確度僅降 0.86 個百分點

2026年9月24日 18:58
站內 AI 整理稿

BottleCap AI has released ThinkingCap-Qwen3.8-27B, the second model in its ThinkingCap series.It is a fine-tune of the Qwen team’s Qwen3.8-27B with one narrow goal: shorter reasoning traces.Across 12 benchmarks, it spends 37.2% fewer thinking tokens on average.Macro-average accuracy moves from 86.

65% to 85.79%, a 0.86pp drop.Deployable?Yes.It drops in for Qwen3.8-27B on vLLM or SGLang, with FP8, NVFP4, GGUF and MLX builds.The repo is gated, and commercial use beyond the small-business license needs a BottleCap agreement.What Problem Does ThinkingCap Target?

Reasoning models often spend more thinking tokens than a question needs.BottleCap’s position is that many of those extra tokens do not change the final answer.The first release in the series applied this idea to Qwen3.6-27B.The objective this time was deliberately conservative.

BottleCap did not try to add knowledge or change answer style.Reasoning ability, instruction following and safety behaviour were meant to pass through untouched.The research team also focused harder on math, reasoning, long-context and agentic benchmarks.

Benchmark Results at xhigh Effort All main numbers use reasoningeffort=xhigh, the chat template default.Every benchmark gets shorter, with cuts ranging from 10.7% to 65.5%.Knowledge and multilingual tasks shrink the most.MMMLU drops 65.5% (1,656 to 571 tokens) and MMLU-Pro drops 57.3%.

GPQA-Diamond falls from 12,772 to 7,267 tokens, a 43.1% cut.IFBench thinks 46.4% less with accuracy nearly flat (79.75% to 79.71%).Long-context retrieval improves.AA-LCR accuracy rises 2.25pp, from 81.75% to 84.00%, with 38.6% fewer thinking tokens.LiveCodeBench v6 edges up 0.07pp while thinking 20.

3% less.Agentic results hold close to the base.τ²-bench gives up 1.01pp for a 30.9% cut.Terminal-Bench 2.1 loses 0.56pp, well inside its ±4.26 interval, for a 10.7% cut.The most expensive trade is AIME 2026.Accuracy falls 3.85pp, from 98.13% to 94.27%, for 30.2% less thinking.

Please note that the 37.2% figure is the mean of the 12 per-benchmark reductions.Pooled mean thinking tokens fall from 15,735 to 12,144.BottleCap also reports a budget curve.Under a 16K-token cap per response, ThinkingCap scores higher than the base model.Truncated traces fall from 0.51% to 0.

34%, and looping from 0.06% to 0.05%.window.addEventListener("message",function(e){var f=document.getElementById("mtp-tcq38-frame");if(f&&e.source===f.contentWindow&&e.data&&e.data.mtpH){f.style.height=e.data.mtpH+"px";}}); How It Interacts With the Effort Dial Qwen3.

8-27B exposes a reasoning-effort setting, and the compression stacks with it.All deltas below compare against the base model at xhigh, averaged over 11 benchmarks.At medium, the base model cuts 52.1% of thinking for -9.16pp.ThinkingCap cuts 60.2% for -9.90pp.At low, the figures are -55.4% and -9.

71pp for the base, versus -62.3% and -10.79pp for ThinkingCap.With thinking off, ThinkingCap trails the base by 5.7pp.BottleCap team recommends xhigh for the best accuracy-to-token balance.It says individual thinking modes will get attention in a future release.

How the Evaluation was Run Both models ran through the same harness on one NVIDIA H200 with vLLM 0.29.0.Sampling was identical: temperature 1.0, topp 0.95, topk 20, minp 0.0.Multi-seed accuracy is the mean with a 95% interval.Seeds range from 32 on AIME 2026 to 1 on MMLU-Pro and MMMLU.

MMMLU uses a fixed 10,000-question sample; the other 11 benchmarks run complete sets.MTP speculative decoding (3 draft tokens) was measured as accuracy-neutral on AIME 2026.It accepted 53% of drafted tokens, about 2.6 tokens per step, matching the base model.

Deployment: Builds, Serving and License The bf16 checkpoint has 28B parameters and accepts image and text input.BottleCap publishes 5 quantized builds: FP8: 31 GB, vLLM, Hopper and Blackwell.NVFP4 weight-only: 21 GB, vLLM, Hopper (Marlin kernel) and Blackwell.NVFP4 W4A4 (AWQ): 23 GB, Blackwell only.

GGUF: 16 to 55 GB, for llama.cpp, LM Studio and Ollama.MLX 4-bit DWQ: 21 GB, for Apple Silicon Macs with 32 GB.Serving uses the base model’s recipe: --reasoning-parser qwen3 with the qwen3_xml tool-call parser on vLLM.Thinking returns in a separate reasoning field.

The license is PolyForm Small Business 1.0.0 plus a BottleCap personal-use grant.Upstream Qwen materials stay under Apache-2.0.Hugging Face lists no inference provider hosting the model yet.Key Takeaways 37.2% fewer thinking tokens on average across 12 benchmarks.Macro accuracy drops 0.

86pp, from 86.65% to 85.79%.AA-LCR long-context accuracy rises 2.25pp; AIME 2026 falls 3.85pp.Drop-in for Qwen3.8-27B: same sampling, same vLLM or SGLang flags.Gated weights under PolyForm Small Business; commercial use needs a license.Check out the technical blog and model weights.

All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost appeared first on MarkTechPost.

Related

相關文章

量子位生成式AI

PCIe顯卡被低估了!內核補齊+通信重構,DeepSeek推理吞吐翻近7倍

是石科技推出Meta-Infer高效推理引擎,透過內核補齊、算子優化與通信重構等軟體優化,於PCIe架構GPU上將DeepSeek-V4.1-Flash輸入吞吐提升近6.87倍,並在國產GPU上也獲得顯著增益。該方法證明無需依賴頂級硬體,靠軟體調校即可縮小與高階GPU的推理效能差距,為長尾硬體提供規模化生產服務的新途徑。

剛剛
鈦媒體生成式AI

梁文鋒用沙盒開啟RSI

字母AI2026.09.24 18:16 · 來自北京全文4232字00:00 / 11:21DeepSeek聯合清華大學發論文,梁文鋒署名文 | 字母AIDeepSeek聯合清華大學發表了一篇文章,署名的最後一人是梁文鋒。文章表面上是說DeepSeek訓練Agent所使用的沙盒,但是論文第6節越看越不對勁。字裡行間寫了三個英文字母:RSI。通過這個沙盒,Agent能創建自己需要的環境,這個環境會再次訓練Agent,被訓練的更強的Agent會創造更好的環境。由此形成了一個小型的RSI閉環。

3 小時前

首部上星AI長劇《後西遊記》幕後:沒有攝影機,100%畫面由Seedance生成,單集成本壓到十幾萬

這部規劃60集、每集約40分鐘的劇集由芒果TV出品、伯璟文化承製,全程沒有一臺攝影機,視頻生成100%交給Seedance完成。上線僅一週,芒果TV正片播放量就衝破1.5億次。把這部劇推出來的,是兩個被傳統影視邏輯卡過的人。總導演李東珅長期拍紀錄片,執導的《河西走廊》豆瓣9.

3 小時前
智東西生成式AI

MiMo-V2.6剛發,小米羅福莉扔出MiMo-V3新架構!引用DeepSeek多項成果

(公眾號:zhidxcom) 作者 | 江宇 編輯 | 李水青 9月24日報道,昨晚,小米MiMo大模型負責人羅福莉發文,提前披露了MiMo-V3的新架構,其核心組件HySparse2率先登場,相關技術論文同步公佈。▲羅福莉發文 簡單來說,這套新架構主要解決Agent越做越多輪之後出現的三個問題:長輸入算得太多、緩存佔得太大,以及如何從越來越長的上下文裡準確找到需要的信息。

5 小時前
鈦媒體生成式AI

OpenAI、Anthropic同日模型大戰,“是兄弟就砍一刀”

光錐智能2026.09.24 16:43 · 來自廣西全文3968字00:00 / 10:28刀刀見骨,AI巨頭為自己畫過的大餅“填窟窿”文|光錐智能,作者 | 魏琳華 ,編輯|劉俊宏9月23日凌晨,OpenAI和Anthropic像約好了一樣,前後腳各自放出新模型:OpenAI端出了GPT-6 Sol和GPT-6 Luna兩款模型,把旗艦GPT-6 Astra的能力蒸進更快更便宜的型號裡;Anthropic則亮出Claude Opus 5.5,性能追平Fable 5.1的同時,成本只有上代Opus 5的六成。

5 小時前