Onton 發布 Ontology 1:神經符號搜尋模型,準確度比全球最佳電商搜尋引擎高出 2.7 倍
重點摘要
總部位於舊金山的搜尋與探索公司 Onton 發布了 Ontology 1,這是一個用於複雜、對話式、多模態商品搜尋的神經符號模型。在由三個獨立 LLM 評審評分的 90 個查詢基準測試中,Ontology 1 的平均 precision@10 達到 0.630,而 Google Shopping 為 0.543,Amazon 為 0.469。且這是在僅索引其目錄約 1% 的情況下達成的。它是否可部署?可以,但不是以可下載的權重形式。Ontology 1 已在 Onton.com 對終端用戶上線,Onton 表示合作夥伴存取會針對在代理式網路上建構的團隊逐案授予。目前沒有公開 API、定價層級或開放檢查點。現階段的採用看起來像是夥伴關係,而非 pip install。適合的企業:中階市場與企業級零售商、市集及代理式商務相關團隊。
Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search.On a 90-query benchmark scored by three independent LLM judges, Ontology 1 reached a mean precision@10 of 0.630, against 0.
543 for Google Shopping and 0.469 for Amazon.It did this while indexing roughly 1% of their catalogs.Is it deployable Yes, but not as weights you download.Ontology 1 is live for end users at Onton.com, and Onton says partner access is granted case by case for teams building on the agentic web.
There is no public API, pricing tier, or open checkpoint for the model itself.Adoption today looks like a partnership, not a pip install.
Company fit: Mid-market and enterprise retailers, marketplaces, and agentic-commerce platforms whose relevance stack already loses on long, requirements-heavy queries.Small catalogs see less benefit, because the failure mode Ontology 1 targets scales with catalog size and listing noise.
Industries: Home decor and furniture today, since that is the only vertical Onton indexes.Onton states the methodology generalizes beyond e-commerce, and that Ontology searches non-product data with essentially no reconfiguration.
Applications: Conversational and multimodal site search, moodboard-driven discovery, negation-heavy filtering, listing and review trust scoring, and grounding layers for shopping agents.
Why keyword and vector retrieval break here Conventional e-commerce assumes intent maps onto categories and attributes: size, price, material, brand.There is no filter for ‘pet-friendly,’ and none for furniture that fits your room.
Onton argues this catalog interface has barely changed in nearly 30 years.Ontology 1 takes a different route.For ‘pet-friendly sectional,’ it does not trust the seller’s label, which may be absent or untrue.
It reasons from properties more likely to be objective — fiber, weave, construction — and flags claims the product data contradicts.It also weighs the source, since some listings game the algorithm and some reviews are bought.
The model builds an explicit, inspectable world model rather than absorbing patterns into weights.When it has no account of ‘pet-friendly,’ it treats that as a gap and works the answer out: cleanability and durability, then polyester upholstery as an indicator.
The learning is reused on later queries such as ‘pet-friendly chair’ or ‘cleanable blue couch,’ and the loop runs continuously.The benchmark: Subtext-Decor-90 Onton released Subtext-Decor-90 with code and data.Three multimodal judges: Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.
5, scored the top 10 visible result cards returned by Onton, Amazon and Google Shopping for each of 90 text queries.P@10 was averaged across judges, with 95% confidence intervals from 10,000 bootstrap resamples.Results: Onton 0.630 [0.571, 0.688], Google Shopping 0.543 [0.490, 0.596], Amazon 0.
469 [0.417, 0.521].Onton won 52 queries outright, Google 19, Amazon 16.Those sum to 87 because Ontology returned fewer than 10 results on three queries, and empty slots were scored as non-relevant.Excluding those slots instead gives Onton 0.665, Google 0.549, Amazon 0.459.
Krippendorff’s alpha across the three judges is 0.465, so absolute P@10 values are noisy and judge-dependent.All three judges still place the engines in the same order.
Image and multimodal queries were excluded from the 90, because Amazon Lens does not support multimodal queries and Google Lens does not return products exclusively.Onton reports a separate 10-query image and multimodal comparison against Google.
Where Ontology 1 loses Failure cases cluster on functional-spec queries where Amazon’s category metadata dominates: ‘lamp that won’t wake my partner if I read at 3am’ (Onton 0.4, Amazon 0.9) and ‘something to put on a weirdly deep windowsill’ (Onton 0.07, Amazon 0.67).
Onton attributes this to catalog breadth and its single-vertical, non-sponsored index, and expects the self-learning loop to narrow the gap.The infrastructure underneath Ontology 1’s knowledge graph runs on Ograph, a custom graph database.
Onton reports one Ograph core beating SuiteSparse:GraphBLAS running on 14 cores, roughly 100× the throughput per core, and a GPU build running 43× faster than the CPU variant, with early runs touching 1000× as the implementation is tuned.
Interactive explainer The embed below walks through the same material in four panels: real Subtext-Decor-90 queries with per-query scores, the pet-friendly reasoning graph drawn step by step, the self-learning loop, and the benchmark chart with confidence intervals and alternate scoring views.
(function(){ window.addEventListener("message", function(e){ var d = e && e.data; if(!d) return; var h = d.ontologyHeight || d.frameHeight; if(!h) return; var f = document.getElementById("ontology1-frame"); if(f) f.style.height = h + "px"; }, false); })(); Key Takeaways Ontology 1 scores P@10 0.
630 on Subtext-Decor-90, ahead of Google Shopping (0.543) and Amazon (0.469).It wins 52 of 90 queries outright while indexing roughly 1% of either competitor’s catalog.The architecture is neurosymbolic: an inspectable knowledge graph that decomposes vague predicates into checkable properties.
Judge reliability is modest (Krippendorff’s alpha 0.465), but all three judges rank the engines identically.Availability is product-first — live on Onton.com, partner access case by case, no open weights or public API.Check out the Technical details and Benchmarks.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the World’s Best E-commerce Search Engines appeared first on MarkTechPost.
Related
相關文章

曝字節訓10億參數大模型,或超Mythos 5,張一鳴、梁汝波先後發聲
字節跳動正在訓練一個參數量高達10萬億的AI模型,規模可能超越Anthropic的Mythos 5。創辦人張一鳴在內部會議中強調編程的關鍵地位,並反對模型蒸餾,認為這只能複製而非超越對手。字節跳動在AI領域持續加大投入,同時在產品端與訓練端採取雙線進攻策略。

AI 需求擠爆雲計算,消息稱 AWS 要求工程師關閉閒置服務器減少資源浪費
因AI需求導致算力緊缺,亞馬遜AWS要求工程師關閉閒置的EC2實例,以減少資源浪費。數據顯示約65%的EC2實例在30天內平均CPU利用率低於20%,AWS因此升級計算優化器自動標記低使用率虛擬機。此外,AWS過去一年新增3.8吉瓦電力容量,仍難以應對GPU雲端實例的龐大需求。

六巨頭定AI插件新標準,撞臉Claude,Anthropic沒上桌
六大科技巨頭(AWS、Anysphere、GitHub、微軟、OpenAI、Vercel)聯合發布AI智能體插件統一開放規範Agent Plugins 1.0.0,旨在統一插件打包格式,減少開發者重複勞動。該規範的結構與Anthropic的Claude Code插件系統高度相似,但Anthropic並未參與制定,而是繼續經營自己的封閉生態。

DeepSeek重啟融資,三年市值對齊騰訊?
DeepSeek重啟第二輪融資,以5000億元人民幣估值尋求籌集80億美元,但網傳一份由小型醫藥私募發起的專項基金募資材料引發網友質疑,後經DeepSeek員工證實部分數據屬實。該公司近期宣布API大幅漲價,可能打破其以低價換規模的估值邏輯,面臨客戶流失風險。市場關注其能否從「價格屠夫」轉型為價值提供商,以及三年內市值能否對齊騰訊等巨頭。

可靈AI核心技術骨幹王鑫濤被曝離職
快手可靈AI核心技術骨幹王鑫濤被曝離職,去向未知,快手官方與本人均未回應。王鑫濤是圖像與視頻生成領域知名開源項目主要作者,被視為可靈從0到1的關鍵推手。其離職發生在可靈完成獨立融資、估值180億美元的關鍵階段,可能影響研發進度與競爭優勢。

AI短劇、漫劇、戀綜、電影、藝人都有了,AI觀眾也不遠了
2026年AI影視內容全面爆發,從短劇、長劇到電影、綜藝,AI製作的作品大量湧現,衛視也開始播出AI短劇。AI演員如方桃子迅速走紅,商業變現能力驚人,廣告報價甚至超過許多真人網紅。AI短劇市場規模已突破220億元,用戶超過6億,但同時也引發了對真人演員就業和內容品質的擔憂。