Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

2026年8月27日 20:05
站內 AI 整理稿

Cohere has released Parse (parse-v5.0), a document parsing model aimed at high-volume enterprise ingestion.It is a 2.3B-parameter vision language model with an 8,192-token context window and a ~4.6GB footprint, built on Cohere Labs’ North-Micro-Vision-Instruct architecture.

Parse takes a PDF, PPT or JPEG page as a base64-encoded data URI and returns Markdown containing text in reading order, tables rendered as HTML, lists, form key-value pairs, image descriptions and bounding box coordinates.There is no separate OCR stage in front of it.

Cohere prices the Parse API at $1.50 per 1,000 pages and positions the model on price-performance rather than peak accuracy — a claim the company supports with a self-reported ParseBench score of 79.2 that, as we detail below, measures three of that benchmark’s five dimensions.Is it deployable?

Yes, in production.Parse is generally available through the Cohere Parse API, Microsoft Foundry, AWS SageMaker, and single-tenant Model Vault.There is no waitlist and no research license.

Which companies: Mid-market teams that already run a RAG stack can start on metered API calls with a free trial key.Large enterprises with residency or air-gap requirements go straight to Model Vault or private deployment.

Seed-stage startups can use it, but the economics only start to matter above roughly 100K pages a month.

Which industries: Cohere targets financial services, insurance, healthcare and life sciences, public sector, telecom, energy and manufacturing — the document-heavy verticals where scanned forms and dense tables are the norm.

Applications: RAG ingestion, intelligent document processing, claims and invoice pipelines, contract and filing search, and giving document context to agents.What is Parse?Parse is a 2.

3B-parameter vision language model built on Cohere Labs’ North-Micro-Vision-Instruct architecture, with an 8,192-token context window and a ~4.6GB footprint.

It accepts PDF, PPT and JPEG pages as base64-encoded data URIs and returns Markdown containing document text, lists, tables rendered as HTML, bounding box coordinates and image descriptions.There is no separate OCR stage in front of it.

The model recovers text and reading order, tables, lists, forms and key-value pairs, images and captions, and the locations of page boundaries and visual elements in one pass.

Nine languages are listed as stable — Arabic, English, French, German, Italian, Japanese, Korean, Portuguese and Spanish — with zero-shot support elsewhere at lower accuracy.Two output modes matter in practice.The default returns a Markdown string per page.

Setting output_format="blocks" returns typed blocks, where a table block carries its HTML, its bounding box and a description.That second mode is what makes citation-level traceability possible.

#mtp-cohere-parse5-explainer hr,#mtp-cohere-parse5-explainer p:empty,#mtp-cohere-parse5-explainer del,#mtp-cohere-parse5-explainer s{display:none!important} #mtp-cohere-parse5-explainer{background:#E8E6DE!important;border:1px solid #D9D4C7!important;border-radius:16px!important;overflow:hidden!

important;margin:26px 0!important} #mtp-cohere-parse5-explainer iframe{width:100%!important;border:0!important;display:block!important;background:#E8E6DE!important} (function(){ var f=document.getElementById('mtp-cp5-frame'); window.addEventListener('message',function(e){ if(e&&e.data&&e.data.

mtpHeight){f.style.height=e.data.mtpHeight+'px';} }); })(); The Benchmark Cohere reports a ParseBench score of 79.2 for Parse, averaged across tables, content faithfulness and semantic formatting, ahead of Mistral OCR 4 (74.5), Azure Document Intelligence (74.3) and Databricks AI Parse (72.4).

ParseBench is a LlamaIndex benchmark of ~2,078 human-verified enterprise pages scored on five dimensions: tables, charts, content faithfulness, semantic formatting and visual grounding.

Cohere’s figure averages three of them and drops charts and visual grounding — the two dimensions where most parsers collapse.Against the public leaderboard, the same vendors score far lower on the full five-dimension overall: Mistral OCR 4 at 60.68, Databricks AI Parse at 60.

68, Azure Document Intelligence (Layout) at 59.64.Azure’s three-dimension average works out to 74.3, which matches Cohere’s figure exactly and confirms the methodology.Cohere Parse is not currently listed on that leaderboard, where LlamaParse Agentic leads at 84.88.So 79.

2 is a vendor-reported subset score, not a leaderboard position.It is a reasonable claim to test on your own documents.What it Costs Cohere prices the Parse API at $1.50 per 1,000 pages.On Model Vault, Parse 5 runs $4.00/hour or $2,500/month for a Medium instance, and $7.

00/hour or $4,300/month for XL.The crossover nobody publishes: at $0.0015 per page, a single Medium instance breaks even at roughly 1.67M pages per month, and XL at roughly 2.87M.Below that, metered API calls are cheaper.

Above it, dedicated capacity wins on price alone — before any argument about data residency, which is usually the real reason enterprises move to Vault.Key Takeaways Cohere shipped parse-v5.0, a 2.3B VLM that converts PDFs, slides and images into Markdown with HTML tables and bounding boxes.

API pricing is $1.50 per 1,000 pages; Model Vault runs $2,500/month (Medium) or $4,300/month (XL).Dedicated capacity only beats metered pricing above roughly 1.67M pages per month.The 79.2 ParseBench figure is vendor-reported across three of five dimensions and omits charts and visual grounding.

Check out the Technical details here and Try it on HF.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown appeared first on MarkTechPost.

Related

相關文章

小米米家掃拖機器人 7C 開售:滾筒活水增壓拖地,國補價 2068 元

作者:浩渺 責編:浩渺 評論: 感謝網友 肖戰割割 的線索投遞! 8 月 27 日消息,今日,小米米家掃拖機器人 7C 現已開售,京東平臺新品上市價 2704 元,疊加 9 折店鋪優惠 + 8.5 折國補,到手價 2068 元。從商品頁面獲悉,這款新品支持滾筒活水拖地,配備恆壓恆溼滾筒拖布,200 轉 / 分鐘高速洗拖,強效清潔咖啡漬、油漬、寵物腳印等頑固汙漬; 配合活水邊拖邊洗,有效減少二次汙染。模擬人手拖地施壓邏輯,滾筒拖布持續穩定貼合地面發力。新品升級 30000Pa 吸力,強效清潔地面灰塵、大小顆粒物、碎屑、殘渣以及地毯纖維深處的寵物毛髮等;配合自動割毛髮防纏主刷、防纏邊刷,清潔時有效減少纏繞問題。AI 攝像頭搭配線激光避障,可精準識別鞋子、吧檯椅、體重秤等 220+ 常見室內障礙物,暗光環境也能靈活應對;結合側面線激光沿邊功能,實現精準貼障清掃。主機可在智能識別邊角環境後,將滾筒外擴伸出,配合雙邊刷深入清掃牆邊、牆角等掃拖死角,有效減少日常清潔盲區。9cm 纖薄機身,輕鬆進入床底、櫃底、沙發底等傳統清潔盲區。主機超聲波傳感器智能識別地毯後,滾筒拖布自動抬升 10mm 並加大吸力,避免打溼地毯的同時,深度清潔隱藏髒汙。主機可穩定跨越 2cm 臺階、門檻等障礙,進入陽臺、廚房、衛生間等高頻清潔場景。主機返回基座後自動洗拖布,拖布清潔完成後啟動熱風烘乾,最快 2 小時烘乾。大功率集塵風機可自動將塵盒內垃圾吸入 2.5L 大容量抗菌塵袋,最多 75 天免倒垃圾。4L 大容量清水箱和 3.5L 大容量汙水箱分離設計,大幅減少頻繁加水換水的麻煩。京東小米米家掃地機器人 7C 水箱版券後 2067.

12 小時前

AI時代的《旅行青蛙》,這款產品正在給出一個新答案

在《旅行青蛙》以輕量放置玩法席捲全球多年之後,AI技術的浪潮正試圖為這類「佛系遊戲」尋找新的可能性。一款被形容為「AI時代的旅行青蛙」的產品近期受到關注,它試圖回答一個核心問題:當遊戲中的角色不再只是依照固定腳本寄回明信片,而是真正擁有自己的「旅程」與「記憶」,玩家與虛擬生命之間的關係會發生什麼改變? 《旅行青蛙》當年的成功,某種程度上來自於它對傳統遊戲機制的反叛。玩家不需要操作青蛙移動、跳躍或完成任務,只需要偶爾收割三葉草、準備便當與道具,青蛙就會自行出門遠行,並在旅途中寄回明信片與特產。

13 小時前

麻省理工學院:AI正在倒逼大學教育重新定位

麻省理工學院近日在一場關於高等教育的研討會上提出一個備受關注的觀點:人工智慧正在倒逼大學重新思考自身定位,而當前教育體系面臨的最大風險,並非學生利用AI作弊,而是這項技術對學習過程本身的侵蝕。這番論述迅速在學術界與科技圈引發熱烈討論,也讓許多人開始重新審視大學教育在AI時代的價值與意義。 長期以來,大學被視為知識傳遞的核心場域,學生透過課堂聽講、閱讀文獻、撰寫作業與考試來累積專業能力。

1 天前

從單點工具到MaaS框架,Gartner 2026曲線為智慧城市AI生態劃出三條賽道

近日,Gartner發佈了最新版的《中國智慧城市和可持續發展技術成熟度曲線》。Gartner研究副總裁相斌斌在接受採訪時指出,今年的曲線已不再將綠色發展與數字化視為兩條平行線,而是將它們融合為驅動新質生產力的同一套KPI體系。相斌斌表示,這是對國家自上而下推動城市韌性建設的回應。Gartner 2026年的這份成熟度曲線清晰地表明,中國智慧城市建設的下半場,將是“集約化、特定化、綠色化”的深度博弈。

1 天前

BOSS直聘第二季度營收23.99億元,AI應用催化增長模式差異化

財報數據顯示,BOSS直聘2026年第二季度取得經營利潤8.63億元,同比上漲32.6%;同期經調整後經營利潤為10.50億元,同比增長19.2%。第二季度,BOSS直聘錄得銷售與營銷費用為5.81億元,較上年同期的4.20億元有所上漲。財報表示,BOSS直聘推進AI應用,在服務用戶規模提升的同時,提升了用戶服務質量,平臺每個求職者平均收穫的達成次數,同環比均實現增長。

1 天前