Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

Most production recommenders are cascades.Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features.Yandex’s Sona Technical Report describes a different design.
Sona is a generative AI model that brings candidate generation and ranking into a single system, replacing the multiple stages typically used in recommendation pipelines.Yandex tested the model in a seven-day live production experiment on its smart speakers.
In an online A/B test, it replaced more than 15 candidate generators, the pre-ranking stage, and the ranking stage with one served transformer.What Problem Does Sona Solve?Cascades split one decision across separately trained models.
Each stage optimizes its own objective, and the ranker only sees what upstream stages let through.Yandex’s previous stack on the Yandex Music surface consumed hundreds of features, including signals from Argus, Yandex’s earlier recommender transformer.
Sona puts candidate generation and ranking around one shared user representation.The encoder reads the listener’s history once per request.A decoder generates candidates.A Ranking Module scores them against the same encoder states.No component uses hand-engineered features.
Inputs are logged event fields (track ID, artist ID, duration, likes, played time, surface flags) and learned Semantic IDs.On Yandex smart speakers, playback can begin without the user first selecting an artist, genre, or mood.The research team describes this as a pure-recommendation setting.
How Sona’s Architecture Works 1.Semantic tokenizer Following the Semantic ID formulation of Rajput et al., every track becomes a tuple of 3 discrete codes.A frozen multimodal LLM reads the mel-spectrogram of the first 90 seconds along with title, artists, and tags.It runs in prefill-only mode.
A 4-layer refinement transformer then aligns those features with listening behavior, using InfoNCE on collaborative track pairs.Residual K-means quantizes the result into 3 codebooks of 32,000 entries each.This beat a CLMR audio baseline: Recall@1000 rose from 0.8111 to 0.8524.2.
Encoder with History Compression Sona attends to 8,192 past events.Full attention over that length is expensive, so the encoder spends depth unevenly.The recent 2,048 events get a 7-layer self-attention stack.Older events pass through cross-attention and 1 full-history layer only.
The paper reports this keeps most of the quality of full attention at about half the inference cost.3.Decoder and Ranking Module A 2-layer decoder emits Semantic ID tuples through constrained beam search with width 1,024.A catalog trie blocks invalid prefixes.
Each tuple expands to every track sharing it.The Ranking Module, which consists of four cross-attention layers, then scores those tracks against the shared encoder memory.Training: A Teacher That Never Ships The Ranking Module learns from a frozen Teacher Ranker.The teacher is a 0.
6B-parameter transformer, also without hand-engineered features.It is trained on a year of engagement events in 2 stages: next-item-prediction pre-training, then multi-head ranking fine-tuning.Removing pre-training dropped weighted pair accuracy from 0.6215 to 0.6153.
The team calls its distillation method Rollout Distillation.During training, the current decoder generates beam candidates.The teacher scores them, together with logged impressions.The Ranking Module regresses onto those scores with mean absolute error.
The joint loss is L = LNTP + Lrollout + L_impression.Both losses update the shared encoder.At serving time, the teacher is removed.Training stays online.Events aggregate into sessions over a 15-minute window, feed a GPU trainer, and new weights reach serving every 10 minutes.
End-to-end latency is 45 minutes at the median and 60 minutes at p99.Serving runs on NVIDIA Triton Inference Server with CUDA graphs and reaches 41% model FLOPs utilization.Results: Online A/B Test on Live Traffic The final experiment ran for 7 days on 15% of randomly selected users per split.
Every change below is statistically significant and relative to the production control: Active Users (primary metric): +4.53% Total Listening Time: +6.30% Likes: +11.42% “Repeat” Commands: +17.99% Deeply Engaged Users: +7.
37% These gains stack on top of improvements retained from earlier deployments.On Active Users, Sona’s uplift is 2.35x the +1.93% increment Argus previously delivered on this surface.Sona vs OneRec vs HSTU: Feature Comparison Sona is not the first end-to-end generative recommender in production.
Kuaishou’s OneRec already serves a single encoder-decoder model.Meta’s HSTU Generative Recommenders reframed recommendation as sequential transduction over user actions in 2024.What Sona combines is a full cascade replacement, no hand-engineered features, and a distilled ranker, validated online.
FeatureSona (Yandex)OneRec (Kuaishou)HSTU GR (Meta)DomainMusic streamingShort videoLarge internet platform, multiple surfacesOne served model replaces the cascadeYes, in A/B testYes, about 25% of total QPSNo, reported as a new architecture for recommendation modelsUser inputsLogged event fields only, no hand-engineered features“Multi-scale feature engineering” pathways, including uid, age, genderUser action sequences (sequential transduction)Item output3-level Semantic IDs, 3 x 32,0003-level Semantic IDs via RQ-KmeansItem IDsRanking signalProduced from frozen 0.
6B Teacher RankerRL with P-Score reward model (ECPO)HSTU ranking modelReinforcement learningNo, fully supervisedYes (ECPO)Not reportedScale reported8,192-event history, 0.6B teacher10x FLOPs of prior ranking model1.5 trillion parametersReported online gain+4.53% Active Users, +6.
30% listening time, +11.42% likes+0.54% and +1.24% App Stay Time+12.4% in online A/B testsPublic code or weightsNoNot in the reportYes, GitHub Sources: Sona, OneRec, HSTU.Online gains come from different platforms and metrics, so they are not directly comparable.
Key Takeaways Sona replaced 15+ generators, pre-ranking, and ranking with 1 served model.No hand-engineered features: only logged events and learned Semantic IDs.A 0.6B teacher trains the ranker, then stays out of serving.A/B test: +4.53% Active Users, 2.35x Argus’s prior gain.
Not deployable externally: no code or weights, not on full traffic.FAQ What is Yandex Sona?Sona is a generative AI model that combines candidate generation and ranking within a single system.
Yandex tested it in a seven-day live production experiment on its smart speakers, where it replaced the existing multi-stage recommendation pipeline for the test group.How is Sona different from OneRec?OneRec uses engineered user feature pathways and RL with a reward model.
Sona uses only logged event fields and distills ranking from a frozen teacher.Check out the Paper for full details.The post Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade appeared first on MarkTechPost.
Related
相關文章
《Gran Turismo 7》迎來首臺中國VGT 同時GT史上30年首次新增電車駕駛教學
新加坡,2026 年 10 月 3 日 —— 今日,在新加坡 Gran Turismo World Series(GT World Series)賽事現場,Gran Turismo系列遊戲製作人山內一典與小米汽車歐洲研發中心設計負責人Jean-Arthur Madelaine現場聯合宣佈:Xiaomi Vision Gran Turismo(小米 Vision GT)即將於10月正式上線《Gran Turismo 7》(GT7),成為該。
A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End
In this tutorial, we implement Kauldron, the JAX training library from Google Research that describes itself as optimized for research velocity and modularity, and we take those two words literally by testing what they actually buy us.
剛剛,iQOO掏出年度旗艦,自研電競芯片性能提升15%,首款電競平板也來了
作者 | 陳駿達 編輯 | 心緣 9月29日報道,剛剛,vivo旗下iQOO品牌發佈了年度旗艦手機iQOO 16,這臺手機搭載了第六代驍龍8超級至尊版,全球首發了由iQOO和三星顯示聯合研發的2K 165Hz三星珠峰屏,並基於iQOO搭建的“3+2遊戲技術版圖”,提升了手機在視效、操控、直播和跨端遊戲等維度的體驗。
抽“錦鯉”享美食!“點亮杭州 碰見好運”服務消費季活動啟動
本文作者: Nemo 2026-09-25 10:17 導語:據瞭解,圍繞“點亮杭州 碰見好運”主題,活動將在9月21日至10月7日期間發放百萬級消費券,覆蓋吃喝玩樂購。9月24日,“點亮杭州 碰見好運”服務消費季活動在西湖區天目裡國際街區正式啟動。
聚焦院外管理提質增效|《急性冠狀動脈綜合徵患者院外長期隨訪管理共識》更新研討,胸痛中心智慧全程管理行動項目正式啟動
近日,第八屆“儒道心學”心血管病學會議、第十屆滬魯心血管病專家論壇、第九屆日照心血管峰會在山東日照召開。由葛均波院士領銜,黃愷、蘇國海、李春潔等數十位心血管領域權威專家參與,會上完成兩大核心動作:一是召開《急性冠狀動脈綜合徵患者院外長期隨訪管理共識》更新研討會,專家集體錨定共識修訂的核心方向;二是胸痛中心智慧全程管理行動項目正式啟動,以專家共識為指引推進先行落地驗證。作為醫療 AI 賦能院外管理創新的先行者,訊飛醫療執行總裁鹿曉亮受邀參會,與學界、業界共同推動心血管院外管理向標準化、智能化、全週期階段邁進。
型別安全 AI Jev 編碼指南:型別決策、校準信心與系統一模型的推測性扇出
在本教學中,我們使用 TypeSafe AI 的第一個系統一模型 Jev,它完全不生成文字:我們向其發送一段程式狀態和一組型別問題,它會回傳選擇、分數和是/否機率,我們的程式碼可以直接根據這些結果進行分支。我們安裝官方 Python SDK,進行第一次呼叫,同時使用三種問題原語,並觀察狀態的形狀如何影響模型能得知的資訊。接著,我們從回傳的機率重新計算已發布的信心統計數據,衡量將十個問題合併為一次呼叫與分開十次呼叫相比的效益,並建立 API 設計所針對的模式:信心門控路由、權重保留在程式碼中的複合評分、型別函式呼叫,以及以模型能處理的方式進行計數。