文檔檢索換路

2026年8月31日 00:00
站內 AI 整理稿

🏆 EVIE-Preview-4.5B Rank #1 on ViDoRe V3 · Rank #1 on ViDoRe V1+V2 The most accurate visual document retriever, with native 128-dimensional token vectors.

🏆 Results • 💾 Index Cost • ⚡ Quick Start • 🔬 Reproducing • 🧠 Architecture • 📚 Citation 🥇 ViDoRe V3 — Rank #1 8 public domains × 6 query languages, nDCG@10.# Model Params Token Dim V3 public 🥇 1 EVIE-Preview-4.5B 4.54B 128D 65.36 🥈 2 webAI-ColVec1.1-8b 8.40B 640D 65.32 🥉 3 webAI-ColVec1.

1-4b 4.54B 640D 63.90 4 nemotron-colembed-vl-8b-v2 8B — 63.54 5 tomoro-colqwen3-embed-8b 8B — 61.60 6 nemotron-colembed-vl-4b-v2 4B — 61.42 7 tomoro-colqwen3-embed-4b 4B — 60.16 8 llama-nemotron-colembed-vl-3b-v2 3B — 59.70 9 colnomic-embed-multimodal-7b 7B — 57.64 10 jina-embeddings-v4 ~3.8B — 57.

54 Two deployment tiers, one checkpoint Visual tokens / page V3 public Vectors / page Raw index / 1M pages (BF16) 768 (training budget) 64.56 751.62 179.2 GiB 1,792 (extrapolated) 65.36 1,763.58 420.5 GiB EVIE was trained at 768 visual tokens per page.

The 1,792 tier is pure test-time extrapolation — the same weights, never trained or fine-tuned at that budget, and never re-exported.That the model does not merely hold up but gains 0.

80 nDCG@10 at more than twice its training budget, improving in 7 of the 8 domains, is a direct read on how well its page representation generalises beyond the resolution it was fit to.Pick whichever tier fits your compute budget.The lighter tier holds a million pages in under 180 GiB.

Per-domain breakdown Model Avg CompSci Energy Finance EN Finance FR HR Industrial Pharma Physics 🥇 EVIE-Preview-4.5B 65.36 80.65 71.36 70.50 54.44 67.34 58.76 69.20 50.62 webAI-ColVec1.1-8b 65.32 80.08 70.12 71.90 54.87 68.55 57.65 67.88 51.50 webAI-ColVec1.1-4b 63.90 80.34 69.50 69.18 53.13 66.

90 56.36 67.25 51.24 nemotron-colembed-vl-8b-v2 63.54 79.30 69.82 67.29 51.54 66.32 56.03 67.19 50.84 tomoro-colqwen3-embed-8b 61.60 75.35 68.41 65.08 49.10 63.98 54.41 66.36 50.13 nemotron-colembed-vl-4b-v2 61.42 78.56 67.48 65.02 49.01 62.39 53.91 66.10 48.86 llama-nemotron-colembed-vl-3b-v2 59.

70 77.09 64.88 64.23 44.41 62.28 51.71 66.04 46.93 colnomic-embed-multimodal-7b 57.64 76.20 63.58 56.57 45.46 58.67 50.13 62.26 48.25 jina-embeddings-v4 57.54 71.81 63.50 59.30 46.10 59.53 50.38 63.09 46.63 EVIE rows measured with reproduce.sh.

Comparison rows are the vendors' published ViDoRe V3 public scores.🥇 ViDoRe V1 + V2 — Rank #1 14 tasks, [email protected] place on the classic boards too.# Model Avg ArxivQA DocVQA InfoVQA ShiftProj SynAI SynEnergy SynGov SynHealth Tabfquad Tatdqa BioMed ESGHL ESG Econ 🥇 1 EVIE-Preview-4.5B 85.77 90.

73 64.53 93.26 93.85 99.63 98.26 98.89 98.89 97.32 81.93 70.17 79.84 64.95 68.53 🥈 2 Ops-Colqwen3-4B 84.90 91.80 66.50 94.00 90.80 99.60 97.30 98.00 99.60 93.60 82.40 65.50 78.60 66.00 64.50 🥉 3 nemotron-colembed-vl-8b-v2 84.80 93.10 68.10 94.60 93.30 100.0 97.90 98.90 99.60 97.70 83.40 66.20 73.

20 60.60 60.80 4 nemotron-colembed-vl-4b-v2 83.90 92.00 67.40 93.30 92.30 99.30 96.20 98.00 98.50 98.10 81.20 64.30 71.40 61.50 60.80 5 colqwen3.5-4.5B-v3 83.70 91.90 66.60 93.60 90.20 100.0 97.10 97.30 98.90 95.90 84.00 65.30 73.80 58.00 59.90 6 llama-nemotron-colembed-vl-3b-v2 83.60 90.40 67.

20 94.70 92.00 100.0 98.00 98.00 98.90 97.30 81.00 63.20 73.10 58.60 58.60 7 tomoro-colqwen3-embed-8b 83.50 91.20 66.40 94.50 87.90 99.30 96.70 97.60 99.10 94.20 80.90 65.50 76.00 60.70 59.50 8 EvoQwen2.5-VL-Retriever-7B-v1 83.40 91.50 65.10 94.10 88.80 99.60 96.60 96.30 98.90 93.60 82.30 65.20 77.

00 59.70 59.10 9 tomoro-colqwen3-embed-4b 83.20 90.60 66.30 94.30 87.40 99.30 96.90 97.20 99.60 94.30 79.90 65.40 74.60 62.40 56.30 10 SauerkrautLM-ColQwen3-8b-v0.1 82.90 93.80 64.70 94.50 90.40 98.60 96.50 96.80 99.30 92.20 84.00 63.30 70.80 57.90 58.00 Tasks 1–10: ViDoRe V1.Tasks 11–14: ViDoRe V2.

Board aggregates: V1 91.73 · V2 70.87.💾 Index Cost Index size is what decides whether multi-vector retrieval actually ships.EVIE emits native 128D token vectors, so the index stays compact at both page budgets.Raw BF16 index 768 tokens/page 1,792 tokens/page 1M pages 179.2 GiB 420.

5 GiB 10M pages 1.8 TB 4.1 TB 1,763.58 vectors/page × 128 dim × 2 bytes × 1,000,000 pages ÷ 2^30 = 420.5 GiB Scoring stays cheap for the same reason: MaxSim is a late-interaction dot product over the token vectors, so a narrower vector cuts the scoring work exactly as it cuts storage.

🧠 Architecture Text Query ───────► ColQwen35 (BiDir Attn) ─────► Query Token Embeddings (128D) │ Late Interaction (MaxSim) ──► Relevance Score │ Document Image ─────► ColQwen35 (Dynamic Vision) ──► Doc Token Embeddings (128D) Vision-Language Backbone — Qwen3.

5-4B with interleaved GatedDeltaNet linear attention and full attention.Compact Projection — contextual token states projected directly into native 128-dimensional representations.Late-Interaction Retrieval — token-level MaxSim between query tokens and document visual tokens.

🌍 Multilingual Queries in English, French, German, Italian, Spanish, Portuguese and Chinese, retrieving over charts, tables, scientific reports, financial filings and scanned forms — including Japanese-language pages.📦 Model Footprint Value Parameters 4.54B Checkpoint (BF16) 8.

5 GB Token embedding 128D Max visual tokens 768 / 1,792 ⚡ Quick Start ColPali Engine is the reference path — every number on this card comes from it.A Sentence Transformers path is also available for late-interaction pipelines already built on that API.Installation pip install "colpali-engine>=0.3.

15" accelerate Or pip install -r requirements.txt if you cloned the repository; that adds pyarrow, which only reproduce.py needs.Python Inference Self-contained — nothing to clone, no local files to prepare.

import torch from huggingfacehub import hfhubdownload from PIL import Image from colpaliengine.models import ColQwen35, ColQwen35Processor modelid = "tencent/EVIE-Preview-4.5B" def enablebidirectionalattention(model): """Encoder-ize the full-attention layers; the GatedDeltaNet layers stay recurrent.

""" for cfg in (model.config, getattr(model.config, "textconfig", None)): if cfg is not None: cfg.iscausal = False for module in model.modules(): if module.class.name in ("Qwen35Attention", "Qwen3Attention"): if hasattr(module, "iscausal"): module.iscausal = False # 1.

Load model and enable bidirectional attention model = ColQwen35.frompretrained( modelid, torchdtype=torch.bfloat16, devicemap="cuda", attnimplementation="flashattention2", ).eval() enablebidirectionalattention(model) # 2.Load processor processor = ColQwen35Processor.frompretrained(modelid) # 3.

Prepare inputs — four example document pages pages = [ hfhubdownload("sentence-transformers/example-documents", f"doc{i}.jpg", repotype="dataset") for i in range(1, 5) ] images = [Image.open(p).convert("RGB") for p in pages] queries = [ "What is the variable represented on the y-axis of the graph?

", "Total outlay is maximum in which year?", ] imagebatch = processor.processimages(images).to(model.device) querybatch = processor.processqueries(queries).to(model.device) # 4.Generate multi-vector embeddings and score with torch.inferencemode(): imageembeddings = model(imagebatch) model.

ropedeltas = None # required before query forward queryembeddings = model(querybatch) scores = processor.score(queryembeddings, imageembeddings) print(scores) # tensor([[17.3750, 10.9375, 7.8750, 7.3438], # [ 6.5938, 13.3750, 6.2188, 6.0938]]) print("Best page per query:", scores.

argmax(dim=1)) # Best page per query: tensor([0, 1]) ⚠️ Apply enablebidirectionalattention(model) once after loading, and reset model.ropedeltas = None before every query forward pass.Both are required to reach the scores above — released colpali-engine (through 0.3.17) builds ColQwen3.

5 with causal masks, which costs about 1.1 on top-hit MaxSim.The same helper ships as bidirectional.py for infer.py and reproduce.py.CLI python infer.py --query "Quarterly revenue report" --image page1.png --image page2.

png Sentence Transformers EVIE also loads as a Sentence Transformers MultiVectorEncoder, exposing the familiar encodequery / encodedocument / similarity API with MaxSim scoring built in.Bidirectional attention is baked into the shipped configuration, so no extra call is needed.

MultiVectorEncoder requires Sentence Transformers 6.0.0, which is not on PyPI yet — install from source until it is released: pip install "sentence-transformers[image] @ git+https://github.com/huggingface/sentence-transformers.

git" from sentencetransformers import MultiVectorEncoder model = MultiVectorEncoder("tencent/EVIE-Preview-4.5B") queries = [ "What is the variable represented on the y-axis of the graph?", "Total outlay is maximum in which year?", ] documents = [ "https://huggingface.

co/datasets/sentence-transformers/example-documents/resolve/main/doc1.jpg", "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc2.jpg", "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc3.jpg", "https://huggingface.

co/datasets/sentence-transformers/example-documents/resolve/main/doc4.jpg", ] queryembeddings = model.encodequery(queries) documentembeddings = model.encodedocument(documents) print(queryembeddings[0].shape, documentembeddings[0].shape) # torch.Size([23, 128]) torch.Size([755, 128]) scores = model.

similarity(queryembeddings, documentembeddings) print(scores) # tensor([[17.3457, 10.8008, 7.8613, 7.3174], # [ 6.5547, 13.3828, 6.2207, 6.0771]]) print("Best page per query:", scores.

argmax(dim=1)) # Best page per query: tensor([0, 1]) Both paths above run on the same four example pages, so they are directly comparable.The two sets of scores agree closely; the small differences come from the attention backend and dtype, and the ranking is identical.

Documents may be file paths, URLs or PIL.Image objects.Text passed to encodedocument is rendered as a query, since this model has no separate text-document format.The default page budget is the 768-token tier.

To score the 1,792-token tier, raise the pixel budget through processorkwargs: model = MultiVectorEncoder( "tencent/EVIE-Preview-4.

5B", modelkwargs={"attnimplementation": "flashattention2", "devicemap": "cuda:0"}, processorkwargs={"size": {"longestedge": 1792 32 32, "shortestedge": 65536}}, ) 🔬 Reproducing Every number on this card is reproducible with the shipped script across all visible GPUs: bash reproduce.

sh On the first run, downloaddata.py fetches the 22 public ViDoRe datasets (~55 GB) from Hugging Face.To reuse an existing directory: bash reproduce.sh /path/to/vidore Target aggregates ViDoRe V1 nDCG@5 91.73 (10 tasks) ViDoRe V2 nDCG@5 70.87 (4 tasks) ViDoRe V1+V2 nDCG@5 85.

77 (14 tasks) ViDoRe V3 public nDCG@10 64.56 (8 domains x 6 languages, 768 visual tokens) ViDoRe V3 public nDCG@10 65.36 (8 domains x 6 languages, 1792 visual tokens) To score the 1,792-token tier directly: python -m torch.distributed.run --nprocpernode=$(nvidia-smi -L | wc -l) reproduce.

py \ --boards v3 --max-visual-tokens 1792 --data-root /path/to/vidore 🎓 Training Details EVIE was trained on approximately 0.8 million high-quality image-query pairs spanning multilingual documents, technical reports, complex financial tables, infographics and document visual QA.

Hard Negative Mining & Evidence Judging Every mined negative is re-judged by a large multimodal judge before it reaches the loss: 🟢 Candidates that actually answer the query are promoted to positives.🟡 Partially relevant or ambiguous candidates are masked out of the loss.

🔴 Only strictly irrelevant pages survive as true hard negatives.Multi-positive rows are group-aware weighted by 1/positivecount so that positives from the same query never penalise each other in-batch.Rows with empty queries, corrupted images or degraded text are dropped.

🙏 Acknowledgements Built on the ColPali Engine by Illuin Technology.Powered by the Qwen3.5-4B vi

Related

相關文章

IT之家模型更新

小米穿戴 8 月更新內容公佈,小米手環 9 等迎多項優化

作者:浩渺 責編:浩渺 評論: 感謝網友 順勢而為 的線索投遞!8 月 31 日消息,今日,小米集團手機部副總裁、可穿戴部總經理張雷分享了小米穿戴 8 月的更新內容。據其介紹,本月多款設備都推送了 OTA 更新,小米手環 9、小米手環 10、REDMI Watch 5、REDMI Watch6 等設備都有不同程度的體驗優化。

剛剛

ChatGPT新模型Bel被曝預訓練完成,參數高達10萬億

24分鐘前奧特曼最後一戰:4個月後,交付AGI23小時前突發,OpenAI徹底斷供Cursor2026-08-29閱讀更多內容,狠戳這裡選靠譜AI,看真實評測查看AI測評官方交流社區加入諮詢項目審核和入駐聯繫項目推薦訂閱號關注下一篇OpenAI 內部,AI 建立了三代「文明」甚至有些感人,但更讓人後怕。

剛剛
IT之家模型更新

消息稱字節跳動豆包大模型 2.2 將推遲發佈,原計劃 8 月推出

作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 Skyraver 的線索投遞!8 月 30 日消息,據白鯨實驗室消息,字節跳動原計劃於 8 月推出的豆包大模型 2.2 將推遲面世。多位接近字節的人士稱,延期的原因是,字節內部希望通過更充分的預訓練和後訓練,把編程、工具調用和 Agent 能力拉上去。

7 小時前

重置額度之神Tibo,劇透了Codex下一次大更新

AI 硬體與軟體之間的競速,在過去幾個月從未如此激烈。就在各家廠商忙著把大型語言模型塞進隨身裝置的同時,OpenAI 旗下程式碼生成工具 Codex 也悄悄進入下一階段的輪替。雖然官方尚未正式公告,但一位被開發者社群暱稱為「重置額度之神」的用戶 Tibo,已經在社交平台上提前揭露了這次更新的部分輪廓,消息迅速在技術圈內擴散,也讓不少依賴 Codex 進行日常開發的工程師開始重新檢視自己的工作流程。

12 小時前
鈦媒體模型更新

誰在給 AI 付賬單:八大科技巨頭的萬億投注版圖

Tech商業2026.08.30 09:34 · 來自上海全文6563字00:00 / 15:49有人當金主,有人造鏟子,有人守入口——誰在為這場萬億AI賭局買單?文 | Tech商業如果把 2026 年的 AI 競賽拍成一張資產負債表,左邊是燒掉的幾千億美元算力、薪資與電力,右邊是尚未完全兌現的收入。

19 小時前