Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

2026年9月29日 04:58
站內 AI 整理稿

Alibaba’s Qwen team has released Qwen-Audio-3.1, a 5-model audio stack spanning ASR, TTS and realtime interaction.The main model is Qwen-Audio-3.1-Realtime, a full-duplex speech model built for voice agents that call tools.

Qwen also cut prices: about 85% on Realtime, about 70% on TTS and up to 95% on ASR.Is it deployable?Yes, as a managed API.qwen-audio-3.1-realtime-plus is live on QwenCloud over WebSocket.No open weights were announced.

What Ships on QwenCloud The model page lists text and audio as both input and output.Context is 262K tokens, with 245K max input and 16K max output.Default limits are 60 requests and 100K tokens per minute.Pricing is $6.4 per 1M audio input tokens and $0.8 per 1M text input tokens.

Text and audio output costs $24 per 1M tokens, with output text not charged.Key features include function calling, web search, structured outputs, context cache and fine-tuning.A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline long-audio transcription.

It supports hot words, speaker separation, punctuation and multilingual plus Chinese dialect recognition.It costs $0.15 input and $0.47 output per 1M tokens.Architecture: 2 Models Behind 1 Voice The system runs 2 models with the same Audio Encoder and LLM design.

A full-duplex decision model predicts whether to keep listening, speak, stop or resume.A speech-to-text model writes the response content as text.A context-aware voice renderer then turns that text into streaming speech.It conditions on conversation history, voice cues and acoustic context.

Training is organized into 3 layers: Think, Act, and Speak and Coordinate.Think: M²-OPD Core-Cocktail SFT re-anchors the audio model to its source text LLM using million-hour-scale paired data.Multimodality OPD follows.

A Text Teacher and a frozen Audio Reference score each token of the student’s own trajectory.This is on-policy distillation, not imitation of pre-written answers.Domain experts for empathy, pragmatic intent and acoustic scenes are then trained with GRPO.

Multi-Teacher OPD merges them into 1 deployable model.Act: Executable Environments Each training domain bundles a tool pool, a stateful JSON database and a natural-language business policy.Domains are seeded from open-source tool and MCP server definitions.

Every task defines 1 of 3 outcomes: a write, a justified refusal, or an unsupported request.Scoring checks terminal state, then permitted writes, then behavioral assertions.A fluent reply cannot rescue a failed state check.GRPO receives rewards at dialogue, milestone and turn level.

Search training penalizes redundant queries with rquery=qmin⁡(1,nrefnpred)r{\text{query}} = q \min \left( 1, \frac{n{\text{ref}}}{n_{\text{pred}}} \right).Mean queries per search call fell from 4.37 to 1.05.Trigger F1 slipped from 60.87% to 58.61%.

Speak and Coordinate This layer decides whether, when and how to speak.On Full-Duplex-Bench v1.5, replies to people talking to someone else fell from 0.13 to 0.03.On v3.0, the filler rate dropped from 0.7590 to 0.2960.There are trade-offs.After interruptions, the unwanted resume rate rose from 0.

035 to 0.130.Interruption stop latency is 1.116 seconds, versus 0.383 for GPT-Realtime-2.Interactive Explainer Explore the Think, Act, Speak loop, duplex decisions, a scored training episode and the search reward.window.addEventListener("message",function(e){if(e.data&&e.data.

mtpQa31H){var f=document.getElementById("mtp-qa31-frame");if(f)f.style.height=e.data.mtpQa31H+"px";}}); Benchmarks at a Glance Against 3.0, Audio MultiChallenge rises from 47.12 to 52.21.The 14-language BBA average climbs from 81.7% to 88.1%.FLEURS WER falls from 9.01 to 3.98.

The τ-Voice figures use a half-duplex speech-to-text adaptation.They are not comparable to official full-duplex results.GPT-Realtime-2 still leads the 50-session human red-team study, 96.00% versus 92.00%.How It Compares FeatureQwen-Audio-3.1-Realtime-PlusOpenAI GPT-Realtime-2Google Gemini 3.

8 LiveInputText, audioText, audio, imageText, images, audio, videoOutputText, audioText, audioText and audioContext / input limit262K128K131,072Max output16K32K65,536Function callingYesYesYes (async by default)Built-in web searchYesNot listedGoogle Search groundingReasoningThinking mode, 2K max reasoningConfigurable effortInterleaved reasoningAudio input / 1M tokens$6.

4$32$3.00Audio output / 1M tokens$24 (text and audio)$64$12.00Open weightsNoNoNoSourceQwenCloudOpenAI docsModel page, pricing Prices are list rates checked September 28, 2026.Token rates are not directly comparable, since each provider tokenizes audio differently.

Key Takeaways Task success on a τ-Voice adaptation rises from 78.4% to 82.0% over 3.0.Replies to background speech drop from 73.0% to 13.0% on Full-Duplex-Bench v1.5.262K context, function calling and web search at $6.4 per 1M audio input tokens.

Tool use is trained with GRPO inside self-evolving executable environments.Multi-turn attack success falls to 26.0% (Chinese) and 23.5% (English).FAQ What is Qwen-Audio-3.1-Realtime?A full-duplex speech model from Alibaba’s Qwen team for voice agents that reason, call tools and manage turn-taking.

Can I self-host it?No open weights were announced.Access is through the QwenCloud API as qwen-audio-3.1-realtime-plus.How much does it cost?$6.4 per 1M audio input tokens and $24 per 1M output tokens for text and audio.Check out the Paper, Qwen-Audio-3.1-ASR, and Qwen-Audio-3.1-Realtime.

All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak appeared first on MarkTechPost.

Related

相關文章

IT之家模型更新

豆包將推個人助理產品:4 月已內測,曾用代號“Spell”

作者:- 責編:汪淼 評論: 9 月 29 日消息,新浪科技獲悉,豆包在個人助理方向正在推進相關產品規劃。一位知情人士表示,延續去年豆包手機助手在 personal agent 方向的策略,今年 4 月,豆包內部就已開始內測個人助理方向的探索產品。

剛剛
鈦媒體模型更新

豆包要切滴滴的蛋糕

字母榜2026.09.29 11:08 · 來自北京全文4286字00:00 / 11:46豆包上線“大出行”入口。文 | 字母榜趕在國慶長假出遊高峰前夕,豆包悄悄上線了“大出行”入口。新入口名叫“出行用豆包”,位於豆包App首頁輸入框的上方,與“對話”“幫我寫作”等標籤並列。點擊進入後,頁面上半部分主要被一鍵規劃出行所佔據,豆包預先寫好了不同的提示詞,比如“晚上想吃口烤鴨,附近哪家正宗”,點擊後即可看到相關建議。

剛剛

千問牽手夸克網盤:授權後翻找整理資料,照片還能秒變海報

用戶在千問裡授權之後,就能直接在對話中翻查、歸置和讀取網盤裡的文件,並據此生成學習工具、工作文檔以及可互動的網頁。從官方說明看,在千問首頁 @ 一下夸克網盤,就能調起夸克 Agent 的技能來"料理"網盤資料。比如忘了某份文件丟在哪個角落,只要用大白話描述要找的東西,千問就會替你去搜;若想幹更復雜的活,也能在工作助理裡啟用"夸克網盤管理"技能,基於盤內資料接著往下做。

13 小時前
雷峰網模型更新

“中國具身大腦”獲國際認可!螞蟻靈波與阿拉伯數字經濟聯盟簽署備忘錄

近日,第五屆全球數字貿易博覽會的重要配套活動之一“中阿數字經濟產業對接會”上,阿拉伯數字經濟聯盟(AFDE)與螞蟻集團旗下具身智能公司螞蟻靈波科技正式簽署框架合作備忘錄。 阿拉伯數字經濟聯盟助理秘書長艾曼·郭內姆(Dr. Ayman Ghoneim)博士與螞蟻靈波科技商業化負責人俞靚代表雙方簽約。根據備忘錄,雙方將建立常態化合作機制,共同推進具身智能與機器人全棧解決方案在阿聯酋、海灣阿拉伯國家合作委員會(GCC)乃至中東市場的部署與本地化落地,並在場景應用、數據模型優化及區域產業生態共建等維度展開全方位合作。

18 小時前

千問牽手夸克網盤:授權後翻找整理資料,照片還能秒變海報

用戶在千問裡授權之後,就能直接在對話中翻查、歸置和讀取網盤裡的文件,並據此生成學習工具、工作文檔以及可互動的網頁。從官方說明看,在千問首頁 @ 一下夸克網盤,就能調起夸克 Agent 的技能來"料理"網盤資料。比如忘了某份文件丟在哪個角落,只要用大白話描述要找的東西,千問就會替你去搜;若想幹更復雜的活,也能在工作助理裡啟用"夸克網盤管理"技能,基於盤內資料接著往下做。

20 小時前5700