Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

2026年9月30日 08:26
站內 AI 整理稿

Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust.It replaces an open-source engine Perplexity had forked for its AI-native search stack.Photon now handles retrieval and ranking for all production traffic.

It also powers a new Fast Search mode in the Perplexity Search API.Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95.Is it deployable?Yes, as a hosted API.Set searchtype: "fast" on POST /search and pay $1 per 1,000 requests.

Photon itself is not open source, so the engine cannot be self-hosted.Why Perplexity Replaced its Old Engine The old engine hit 3 limits as the index grew: Tail latency: Production p99 sat near 800 ms.The dataset exceeded RAM, so mlock was not an option.

Cold reads triggered major page faults that stalled queries.Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes.Slow recovery: Deploying and syncing an extra cluster could take more than a week.Recovery also raised the share of partial responses.

Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.How Photon Works A load balancer routes each request to a Photon broker.The broker fans out to a shard group and watches for timeouts.

Each shard runs retrieval, initial ranking, and second-stage ranking.The broker then merges candidates and fetches key document fields.Adaptive posting lists: Short lists sit inline within a single page.Longer lists split into blocks of fixed document ID ranges.

Sparse blocks store sorted offset arrays and use galloping search.Dense blocks use bitmaps, so membership becomes a single bit lookup.Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists.Cheap presence checks bound each candidate’s maximum score first.

Exact term frequencies are read only when a candidate can clear the threshold.Docblob records: Each document gets a compact record of frequencies, field masks, and positions.Terms use Elias-Fano encoding, so ranking decodes only the matched terms.Ranking a candidate needs just 1 lookup per document.

Batched async reads: Record offsets are known upfront, so disk reads go out in batches through iouring.The cache checks the whole batch first.Readers take no locks, and eviction uses CLOCK instead of a shared LRU list.

Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes.A controller rotates serving groups one at a time and warms caches with replayed search-log queries.A full web index now builds in a single-digit number of hours.

Interactive Explainer: Inside Photon window.addEventListener("message",function(e){if(e.data&&e.data.mtpEmbed==="px-photon"&&e.data.height){var f=document.getElementById("mtp-photon-frame");if(f)f.style.height=e.data.

height+"px";}}); Production Results p99 retrieval and ranking latency fell from about 800 ms to about 65 ms.This covers Photon’s stages only.Photon runs on about 20% fewer serving machines than the old content nodes.It stores about 2.

5x as much data per document, which Perplexity used to improve ranking quality.Pinning the same dataset with mlock would need an estimated 4.6x the resident memory Photon uses today.Index version switches no longer cause latency spikes.

Fast Search: Speed and Cost for Agents Fast Search pairs Photon with lighter ranking tuned for agentic workflows.Perplexity tested it on 6 benchmarks: WideSearch, BrowseComp, DSQA, FRAMES, SEAL-0, and SEAL-Hard.Across 3,554 tasks, Fast scored 64.3% at $59.73 in estimated model-plus-search cost.

The default preset scored 64.0% at $187.60, so Fast was about 68% cheaper.The trade-off shows up in broader search quality.On internal long-tail benchmarks, relevance (DCG) fell from 2.45 to 2.21.Answer availability dropped from 0.596 to 0.567, a loss of 2.9 percentage points.

Perplexity recommends Fast for day-to-day agent loops and the default for hard, ambiguous queries.Copy CodeCopiedUse a different Browsercurl -X POST 'https://api.perplexity.

ai/search' \ -H "Authorization: Bearer $PERPLEXITYAPIKEY" \ -H 'Content-Type: application/json' \ -d '{"query": "latest stable Rust release", "searchtype": "fast", "maxresults": 5}' On Python SDK 0.43.4 and 0.43.5, pass extrabody={"searchtype": "fast"} per the docs.

Fast Search vs Closest Competitors FeaturePerplexity Fast SearchExa InstantParallel Search TurboTavily ultra-fastRequest parametersearchtype: "fast"type: "instant"mode: "turbo"searchdepth: "ultra-fast"Vendor-reported latency160 ms p50, 230 ms p95 (blog)~250 ms typical (docs); sub-200 ms at launch~200 ms (docs)No figure published; lowest-latency depth (docs)List price per 1K requests$1 (pricing)$4 for up to 10 results (pricing)$1 (docs)1 credit: $8 pay-as-you-go, $5 to $7.

50 on plans (credits)Results per request1 to 2010 in base price, $1 per 1K per extra resultNot specifiedNot specifiedKnown limitsLower relevance than default presetExtra results billed separatelyEnglish and Japanese queries onlyLower relevance than other depthsLaunchedSep 24, 2026Feb 12, 2026Jul 13, 2026 (blog)Jan 5, 2026 (blog) All latency figures are vendor-reported under different setups, so they are not like-for-like.

Key Takeaways Photon is Perplexity’s Rust retrieval and ranking engine, now serving all production traffic.Production p99 latency dropped from about 800 ms to about 65 ms.Fast Search reports 160 ms p50 and 230 ms p95 at $1 per 1,000 requests.

Fast cut estimated agent task cost by about 68% at comparable task quality.It trades some retrieval relevance, so keep the default preset for hard queries.Check out the technical details and Fast Search docs.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms appeared first on MarkTechPost.

Related

相關文章

IT之家其他AI

特朗普簽署行政令把 AI 改叫 SI,推出美國政府對話式問答網站

作者:汪淼 責編:汪淼 評論: 感謝網友 火瓶座、有鯽雪狐、不一樣的體驗、軟媒新友2547452、呼呼啦啦啦噠 的線索投遞!9 月 30 日消息,美東時間週二(9 月 29 日),美國總統特朗普正式簽署行政令,下令所有行政部門和機構使用“超級智能”(Super Intelligence,簡稱 SI)一詞代替人工智能(AI)。

剛剛
雷峰網其他AI

現在不買帶線控底盤的車,三年後註定會後悔

如果要找出2026年車市升溫最快的技術賽道,線控底盤一定算一個。此前在2024年,線控轉向首次在量產車上出現,當時行業裡的共識是:好東西,但離普通人太遠。2026年7月1日,由上汽集團牽頭制定的線控轉向國家標準《汽車轉向系基本要求》(GB17675-2025)正式實施,政策閘門打開。2026年9月16日,理想i9上市,線控轉向加後輪轉向標配。李想專門在微博上談起這兩項配置,他的思考是,大空間與好操控,需要靠技術進步同時實現,不能一直依賴把車做大。線控轉向和後輪轉向,正是理想為改善體驗投入的方向。由此可見,普及線控底盤,正在成為行業共識。9月23日,全新一代智己LS6上市,19.79萬元起,全系標配全線控底盤,上市57分鐘新增訂單突破11000臺。不到兩年,一項技術完成了從買不起到買得到、再到不買就落後的三次下探。在很多人還將它視為旗艦車型的稀缺配置時,智己直接把它拉到了20萬。一位產品經理對雷峰網這樣形容如今線控底盤這項配置:今天買一臺不帶線控底盤的車,和2020年買一臺不支持整車OTA的車、2022年買一臺沒有激光雷達的車、2023年買一臺還在用400V的車、2025年買一臺沒有數字底盤的車,是同一個道理,當時都覺得夠用,三年後準後悔。01為什麼現在不買,以後就會後悔線控底盤拆掉的,是那根把方向盤和車輪“焊死”的轉向柱。傳統機械轉向裡,方向盤和車輪之間是硬連接。人機共駕時,系統要打方向,就會跟人手較勁,這就是很多人吐槽過的搶方向盤。而L3以上要求駕駛員脫手,機械連接反而成了障礙。線控把那根柱子拆了,轉向變成電信號,智駕系統才可能真正全權接管、做全冗餘。這正是不買就後悔的癥結所在,因為OTA改的是軟件,而線控改的是硬件。今天買一臺不帶線控的車,三年後軟件再怎麼升級,手腳也跟不上大腦。這是後期補不回來的硬件能力,不像車機卡了可以等一次更新,底盤缺的東西,從提車那天起就定型了。

5 天前
雷峰網其他AI

ECCV 2026 開幕:李飛飛團隊獲時間檢驗獎,7000人擠爆馬爾默

本文作者: 叢末 2026-09-23 09:51 導語:93場Workshop,吳佳俊成全場最忙碌的人!93場Workshop,吳佳俊成全場最忙碌的人!作者丨幸麗娟 編輯丨岑 峰 瑞典當地時間8點,ECCV 2026 正式迎來開幕。清晨,馬爾默全城各區的參會者湧向會場,還有一群人乘坐火車橫跨厄勒海峽大橋,從哥本哈根趕來。

1 週前
IT之家其他AI

逾七成美國人擔憂 AI 公司防控不力,超半數支持放緩開發

作者:簫雨 責編:遠洋 評論: 北京時間 9 月 22 日,據路透社報道,就在人們越來越擔憂 AI 技術的未來走向之際,路透社 / 益普索的最新民調顯示,接近四分之三的美國人 (大約 73%) 擔心 AI 公司在防止 AI 對社會造成嚴重危害方面做得還不夠。

1 週前
IT之家其他AI

美聯邦航空管理局斥資 8.75 億美元打造 AI 空管系統,治理空中交通問題

作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞!9 月 19 日消息,據《華爾街日報》當地時間 17 日報道,美國聯邦航空官員將在全美最繁忙的航空走廊之一啟用一套 AI 空中交通管理工具,最快當地時間 21 日上線。據美國政府和航空業官員透露,這套名為 Smart 的軟件將率先在華盛頓特區一帶投入使用,隨後逐步推廣至全美。

1 週前