The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T

2026年10月5日 04:10
The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T
站內 AI 整理稿

In April 2023, Alibaba Cloud demoed a chatbot whose name roughly means ‘truth from a thousand questions.’ Three and a half years later, its descendant ships open weights with 2.4 trillion parameters.This is the story of how Qwen got there, release by release.window.

addEventListener('message',function(e){if(e.data&&e.data.mtpQwenTl&&e.data.h){var f=document.getElementById('mtp-qwen-tl-frame');if(f&&e.source===f.contentWindow){f.style.height=e.data.h+'px';}}}); Chapter 1 — 2023: a thousand questions Alibaba moved after ChatGPT, but not by much.

On April 7, 2023, Alibaba Cloud began handing invitation codes to corporate customers for a model called Tongyi Qianwen.The name draws partly on the philosopher Mencius.Four days later, at the Alibaba Cloud Summit in Beijing, then-CEO Daniel Zhang unveiled it publicly.

Alibaba said it would roll the model into every business, starting with DingTalk and the Tmall Genie voice assistant (China Daily).The real turn came in August.On August 3, 2023, Alibaba open-sourced Qwen-7B and Qwen-7B-Chat, a direct answer to Meta’s Llama 2.Qwen-7B was pretrained on over 2.

2 trillion tokens with a 2,048-token context.Its license was free for commercial use below 100 million monthly users.Vision followed within weeks.Qwen-VL, the first vision-language branch, launched in late August 2023.

On September 13, 2023, Tongyi Qianwen opened to the general public, a sign of Chinese regulatory approval.The same month, the team published the Qwen Technical Report on arXiv.The year closed with scale.Alibaba released its 72B and 1.8B models for download around December 1, 2023.

Qwen now spanned laptop-sized to frontier-sized open weights.Chapter 2 — 2024: the open-weight machine 2024 is the year Qwen became a default choice for developers.The team shipped 3 generations in 8 months.Qwen1.5 (February): On February 5, 2024, the team released Qwen1.

5, framed as a beta of Qwen2.It covered 0.5B to 72B dense models with stable 32K context at every size.A 14B MoE with 2.7B active parameters followed.So did a 110B dense model, the family’s first above 100B.Qwen2 (June): The team announced Qwen2 in early June 2024 in 5 sizes, from 0.5B to 72B.

That lineup included Qwen2-57B-A14B, its first open mixture-of-experts flagship.Training data added 27 languages beyond English and Chinese.The 7B and 72B instruct models handled up to 128K tokens.The licensing shift mattered most: every size except 72B moved to Apache 2.0.

Specialists arrived over the summer.Qwen2-Math launched in August, along with Qwen2-Audio.Qwen2-VL followed at the end of August, able to analyze videos over 20 minutes long.Qwen2.5 (September): At the Apsara Conference on September 19, 2024, Alibaba released over 100 open-source models at once.

The Qwen2.5 Technical Report says pretraining data grew from 7 trillion to 18 trillion tokens.Sizes ran 0.5B to 72B, with 128K context and 8K-token generation.Adoption was already real.

Alibaba said Qwen models had passed 40 million downloads and inspired over 50,000 derivative models on Hugging Face.Coding and reasoning (November–December): Qwen2.5-Coder shipped its full family on November 11, 2024.Then came the first reasoning model.

QwQ-32B-Preview arrived in late November under Apache 2.0, as an open challenger to OpenAI’s o1.QVQ-72B-Preview, an experimental visual reasoning model, closed the year on December 24.Chapter 3 — Early 2025: answering DeepSeek January 2025 belonged to DeepSeek-R1.

Qwen’s response came in weeks, not months.On January 26, the team shipped Qwen2.5-VL in 3B, 7B and 72B sizes.TechCrunch noted it could control PCs and phones.Three days later, on the first day of Lunar New Year, Alibaba launched Qwen2.5-Max, a large-scale MoE model.

Reuters reported Alibaba’s claim that it surpassed DeepSeek-V3.The bigger statement came on March 6.QwQ-32B, built on Qwen2.5-32B and trained with reinforcement learning, shipped under Apache 2.0.Qwen claimed performance comparable to DeepSeek-R1, a 671B model.

VentureBeat put the hardware gap at about 24 GB of VRAM versus over 1,500 GB.Alibaba’s Hong Kong shares rose more than 7% that day.Multimodality kept pace.Qwen2.5-Omni-7B arrived on March 26 under Apache 2.0.It took text, images, audio and video as input, and answered in text or speech.

CNBC framed it as a model for cost-effective AI agents.Chapter 4 — 2025: Qwen3 and the trillion-parameter line Qwen3 (April): On April 29, 2025 (Beijing time), the team released Qwen3: 6 dense models from 0.6B to 32B and 2 MoE models.The flagship was Qwen3-235B-A22B, with 22B active parameters.

Every model shipped under Apache 2.0.The key feature was hybrid thinking.One model could reason step by step or answer instantly, toggled by the user.The Qwen3 Technical Report lists pretraining on about 36 trillion tokens.Language coverage jumped from 29 to 119 languages and dialects.

TechCrunch called it a family of “hybrid” reasoning models.The 2507 refresh (July): On July 21, 2025, Qwen split hybrid thinking back apart.Qwen3-235B-A22B-Instruct-2507 shipped as a non-thinking model with a 262K native context.Separate Thinking-2507 checkpoints followed (Hugging Face collection).

Agents and images (July–August): Qwen3-Coder-480B-A35B launched on July 22, 2025 for agentic coding.Alibaba paired it with Qwen Code, an open-source terminal coding agent (Computerworld).On August 4, the team open-sourced Qwen-Image, a 20B MMDiT model built for text rendering inside images.

An editing variant, Qwen-Image-Edit, was open-sourced on August 18, 2025.The September sprint: Qwen3-Max-Preview, the first Qwen model over 1 trillion parameters, appeared on September 5.Unlike past flagships, it shipped API-only.Days later came Qwen3-Next-80B-A3B, an architecture bet.

It mixed Gated DeltaNet linear attention with gated attention in a 3:1 layout.The model card claims 10% of Qwen3-32B’s training cost and 10x inference throughput past 32K tokens.On September 22, Qwen3-Omni and Qwen3-VL followed.At Apsara on September 24, Alibaba formally launched Qwen3-Max.

CEO Eddie Wu said spending would exceed the RMB 380 billion (US$53 billion) three-year AI and cloud plan.By year’s end, Qwen was national infrastructure for others.In November 2025, AI Singapore said its Sea-Lion model would switch from Llama to Qwen as its base.Qwen also became a consumer brand.

The Qwen App entered public beta on November 17, 2025.Alibaba said it passed 10 million downloads in its first week.Chapter 5 — 2026: 4 generations in 6 months A fast start (January–February): On January 26, 2026, Qwen released Qwen3-Max-Thinking, a closed reasoning model.

Qwen said it was comparable to GPT-5.2-Thinking, Claude Opus 4.5 and Gemini 3 Pro across 19 benchmarks (TestingCatalog).Qwen3-Coder-Next, a small hybrid model for agentic coding, followed on February 2.Qwen-Image-2.0 merged generation and editing into one model on February 10.Qwen3.

5 (February): On February 16, 2026, Lunar New Year’s Eve, Alibaba unveiled Qwen3.5 for what it called the agentic AI era.The open Qwen3.5-397B-A17B activates 17B of 397B parameters.It runs on the hybrid linear-attention design that Qwen3-Next had previewed.A hosted Qwen3.

5-Plus added a 1M-token context.A medium series followed on February 24.Qwen said its Qwen3.5-35B-A3B beat the older Qwen3-235B-A22B-2507.Small on-device models arrived on March 2.A leadership break (March): On March 3, tech lead Junyang Lin posted that he was stepping down.

Reuters reported he was the third senior Qwen figure to leave in 2026.Post-training head Yu Bowen also left, and coding lead Hui Binyuan had joined Meta in January (Yicai).The trigger, per 36Kr, was a plan to split Qwen into horizontal pre-training, post-training and modality teams.

On March 16, Alibaba put its AI units under a new group led by CEO Eddie Wu.By then, Alibaba had released more than 400 open Qwen models, downloaded over 1 billion times.Qwen3.6 (April): The cadence did not

Related

相關文章

量子位生成式AI

限時28天!OpenAI承諾沒新功能就重置,網友:只想要Opus

有改進就體驗,沒改進就重置,橫豎不虧。henry 發自 凹非寺 | 公眾號 QbitAI 要我說,OpenAI(SI)已經瘋掉了!專挑在放假的時候整活~ 剛剛,賽博義父、OpenAI Codex負責人Tibo放話:接下來的28天,每天都要二選一: 要麼交付一項對大多數Codex/Work用戶明顯有用的改進,要麼來一次完整的額度重置。

剛剛
IT之家生成式AI

軟銀集團孫正義罕見發出 AI 安全警告,呼籲各國攜手應對威脅

作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 烏蠅哥的左手、不一樣的體驗 的線索投遞!10 月 5 日消息,據彭博社報道,軟銀集團創始人孫正義一直被視為人工智能最堅定的擁護者之一。然而他近日坦言,隨著 AI 能力的突飛猛進,就連他也對伴隨而來的安全風險深感擔憂。

剛剛
IT之家生成式AI

Meta 的 AI 助手 Muse 被曝可深度分析用戶社交關係,併為每位聯繫人建立個人檔案

作者:遠洋 責編:遠洋 評論: 10 月 5 日消息,Meta 的全新個人助手 Muse 已經迅速走紅,數百萬用戶下載了這款人工智能智能體,把它和自己的銀行賬戶、消息軟件或者健康數據進行綁定,交由它代為完成各項事務。不過就在 Muse 在普通消費者群體當中迅速普及之際,來自應用內部的數據,也讓外界得以窺見這款產品整理、向用戶展示信息的運行邏輯。

剛剛
IT之家生成式AI

OpenAI 奧爾特曼:人工智能的巨大效益值得承擔部分風險

作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 不一樣的體驗 的線索投遞!10 月 5 日消息,OpenAI 首席執行官薩姆 · 奧爾特曼(Sam Altman)近日在接受《Decoded》欄目專訪時表示,他的公司與競爭對手 Anthropic 在人工智能監管問題上,仍然存在根本性的世界觀分歧。

剛剛
鈦媒體生成式AI

9秒刪庫,700個AI失控,豪擲513億買不來安全?

新質動能2026.10.05 08:50 · 來自北京全文4144字00:00 / 12:04AI從工具變成同事,安全就從選項變成前提。文 | 新質動能過去幾年,企業談AI安全,大多數時候指的是內容安全:過濾有害輸出、防止數據洩露、滿足合規備案。

剛剛