Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

2026年10月2日 15:23
Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
站內 AI 整理稿

Datalab has released OmniExtractBench, an open benchmark for structured document extraction.It tests how accurately a system fills a JSON schema from a PDF.The benchmark pools 620 documents from 4 existing benchmarks.One deterministic scorer grades all of them and explains each decision.

The release lands while extraction vendors publish their own leaderboards.Datalab argues those leaderboards are hard to compare or audit.OmniExtractBench is its attempt at a shared yardstick.Is it deployable?Yes, the scorer installs from PyPI as omni-extract-bench (v0.1.7, Python 3.

11+, SciPy only) under Apache 2.0.Rerunning vendors requires your own API keys and paid credits.What is OmniExtractBench?OmniExtractBench is a structured extraction benchmark built by Datalab.Each task gives a system a PDF and a JSON schema.

The system returns JSON, which is scored value by value against a gold file.The code is on GitHub, and the data is on Hugging Face under CC BY 4.0.

The 4 flaws it targets Datalab’s launch post names 4 recurring problems with existing extraction benchmarks: Bias: documents and scoring can favor the vendor that built the benchmark.Opaque harnesses: a low score may reflect a broken harness, not a weak model.

Unclear scoring: readers cannot tell why a given document scored low.Narrow document variety: some suites hold only dense tables, others only clean, unscanned files.

Where the 620 documents come from SuiteDocumentsUpstream publisherContentExtractBench329LlamaIndexForms, filings, decksInternal202Datalab (synthetic)Dense scalar schemas, small documentsLongExtractBench47micro1 (commissioned by Reducto)Very large tablesLongArray-Extract42ExtendLarge tables with repeated scalars Regulatory filing forms are the largest category, at 88 documents.

128 documents are a single page.At the other end, 33 documents over 100 pages hold 40% of all pages.Datalab’s own synthetic suite is the second largest share.How the scorer works The scorer flattens prediction and gold JSON into addresses, which are paths to single values.

It normalizes each value first, so “03/31/2024” matches “2024-03-31”.Tables are the hard part.Compared by position, one missed row shifts every row after it.We reran the scorer on a 100-row table missing its first row.Positional comparison scored 0%, while OmniExtractBench scored 99%.

The fix is content-based pairing with the Hungarian algorithm.ExtractBench and LongArray-Extract already align rows this way.OmniExtractBench adds a verdict layer on top.6 verdicts per value matched: paired, and the values agree.misread: paired, but the values differ.

unfound: gold has a value, the prediction does not.fabricated: the schema allows it, gold is silent, the prediction fills it.inventeditem: part of a predicted row that pairs with nothing.inventedfield: an address the schema never declared.Accuracy is matched values over all verdicts.

Precision divides matched values by predicted values.Recall divides them by gold values.The full rules are in the metric spec.The null rule Empty strings, None and whitespace count as omissions, so those addresses are dropped.Strings like “NA” or “-” remain real answers.

This blocks a quiet exploit: padding a schema with empty optional fields to earn free matches.In our test, padded null fields added 0 verdicts.(function(){window.addEventListener("message",function(e){if(!e.data||e.data.oebFrame!=="oeb1")return;var f=document.

getElementById("mtp-oeb1-frame");if(f&&e.source===f.contentWindow){f.style.height=e.data.

h+"px"}});})(); How it compares with other extraction benchmarks BenchmarkPublisherDocumentsRow alignmentPer-value explanationScorer licenseData licenseOmniExtractBenchDatalab620, from 4 sourcesHungarian, by contentYes, 6 verdict typesApache 2.0CC BY 4.

0ExtractBenchLlamaIndex370HungarianPer-field diffs, HTML reportApache 2.0Apache 2.0LongArray-ExtractExtend45, syntheticHungarianPer-document scoreNot statedCC BY 4.0LongExtractBenchmicro1225, of which 50 publicBy row keyNot statedMITCC BY 4.0, labels only Sources linked on each benchmark name.

Checked September 27, 2026.Who misses fields and who invents values Datalab scored 10 system configurations on the full corpus.Its accurate mode led at 93.85 accuracy.Datalab balanced (93.48) and Reducto deep_extract v2 (93.47) are effectively tied.

Precision and recall then show how each system fails.Balanced: Datalab (both modes) and Reducto keep precision and recall within 0.6 points.Leans to misses: GPT 5.6-sol posts 95.11 precision but 84.99 recall.It loses 11.88% to unfound values.Gemini and Claude show the same pattern, less sharply.

Leans to invented values: LlamaExtract has 93.13 recall but 86.57 precision, losing 9.03% to fabricated values.Extend loses 4.01% to invented items.Low on both: Mistral OCR 4.1 and Azure Content Understanding trail on both metrics, with recall lower still.

How each system falls short of 100%, by verdict type.Source: Datalab.Run it yourself Copy CodeCopiedUse a different Browseruv pip install omni-extract-bench oeb score --pred pred.json --gt gold.json --schema schema.json --verdicts Install the [benchmark] extra and run oeb benchmark to rerun vendors.

Runs are resumable, and each provider needs its own credentials.Datalab also suggests testing its playground on your own documents.Key Takeaways OmniExtractBench pools 620 documents from LlamaIndex, micro1, Extend and Datalab suites.A deterministic scorer gives every value 1 of 6 auditable verdicts.

Content-based row pairing stops one missed row from zeroing a table.Dropping null addresses stops schema padding from inflating scores.Datalab accurate leads at 93.85; Datalab balanced and Reducto tie near 93.5.FAQ What does OmniExtractBench measure?

It measures how accurately a system extracts values from a PDF into a JSON schema, scored per value.Is OmniExtractBench open source?Yes.The scorer is Apache 2.0 on GitHub and PyPI.The dataset is CC BY 4.0 on Hugging Face.Who built OmniExtractBench?Datalab built it.

Datalab also maintains the open-source Marker and Surya document tools.Thanks to the Datalab team for the resources behind this article.Datalab supported and sponsored this content.

The post Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks appeared first on MarkTechPost.

Related

相關文章

具身智能不只在地上跑:給無人機裝上機械臂,它們要去天上「擰螺栓」了 | IROS 2026

本文作者: 張豈萍 2026-09-30 14:29 導語:具身智能的研究邊界,正從地面移動操作向空中協同作業延伸。具身智能的研究邊界,正從地面移動操作向空中協同作業延伸。作者丨張豈萍 編輯丨幸麗娟 機械臂抓取、移動機器人導航和多機器人協作,是具身智能研究的常見切口,但大多仍延續“地面機器人”範式:身體默認是地面移動機器人,環境是結構化可通行空間與離散物體,任務是感知、移動、抓取、搬運等。

2 天前

美議員提出重磅法案:政府出臺安全防護機制前,AI 不得進行遞歸式自我改進

作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞!9 月 30 日消息,當地時間 28 日,據美國 CNBC 報道,代表硅谷選區的民主黨聯邦眾議員羅 · 康納準備提出一項 AI 監管法案,引入嚴格責任標準,並禁止 AI 在缺乏政府安全防護機制的情況下進行遞歸式自我改進。

2 天前
MarkTechPost AI研究與前沿

NVIDIA 研究人員推出 Physis-Lang:自我進化的物理語言,在物理基準測試上讓 Cosmos 3 超越 Veo 3.1

影片世界模型能渲染出令人信服的片段,但仍可能違反物理定律。奶油像油漆般擴散,球穿牆而過。來自 NVIDIA、MIT 和牛津大學的團隊主張,修正之道在於語言本身,而非額外的視覺、潛在或數值訊號。他們的框架 Physis-Lang 將物理語言視為一種共享且可最佳化的表徵,同一份文字可用於資料整理、模型訓練與推論。

2 天前

機器人“拜師”蘇繡!最細的活,最硬的考題

機器人前瞻(公眾號:robot_pro) 作者 | 劉俐杉 編輯 | 許麗思 聽說機器人甚至已經學會刺繡了? 如果讓機器人搬運重物、抓取零件,你可能早已見怪不怪,讓它搭積木、夾豆子、擰瓶蓋,也已經不算什麼新鮮事了。 但今天上午,臨界點發布的一段視頻,讓靈巧手的“手上功夫”又往前了一步。 鏡頭裡,靈巧手捏著一枚繡花針,將細線穿過針孔,隨著針尖的起落,手指在絲線中穿梭。直至鏡頭拉近,你還會看到靈巧手面對一幅蘇繡,與蘇繡老師配合完成合繡。更加值得注意的是,整段視頻均為實拍,沒有加入任何CG動畫。

2 天前