Harvey Introduces Harvey Tenet: A Kimi K3 Base Post-Trained with Fireworks for Long-Horizon Legal Agent Work
Harvey has released Harvey Tenet, its first post-trained model, as a research preview as of today.Tenet is a Kimi K3 base post-trained with Fireworks through asynchronous reinforcement learning on long-horizon legal work.
The training corpus combined synthetic data, publicly available legal data, and human expert data.Harvey states no customer data was used.
Against the base K3 model, Tenet completes almost twice as many held-out tasks on Harvey’s Legal Agent Benchmark (LAB) and 20% more on LAB: Contracts, raising all-pass rate by 9 and 2 percentage points respectively.Harvey reports state-of-the-art on LAB: Contracts and second place on LAB.
The gains also transferred, untrained, to Mercor’s APEX Agents and Crosby’s Redline Bench.The stated goal is twofold: build frontier legal intelligence on open-weight models, and give law firms a path to own their own specialized models.Is it deployable?
Not yet, Harvey Tenet is a research preview announced on August 20, 2026.Harvey has not published weights, a model card, or an API endpoint.
The base model is open-weight; Tenet itself is Harvey’s own checkpoint, and the company says the work will move “from research to production” inside Harvey’s products over time.What ships today is the recipe, not the artifact.Company tier: Enterprise only.
Access runs through Harvey’s platform, which is sold to law firms, mid-sized firms, and in-house legal teams.A lab with an RL stack could reproduce the method; training used roughly 150 NVIDIA B300 GPUs over two months.
Industries: Legal services, corporate in-house legal, private equity and investment banking (M&A diligence), plus regulated sectors where contract volume drives cost — insurance, financial services, healthcare, energy.
Applications: M&A due diligence memos over datarooms, contract drafting, review and redlining, structured extraction across up to 10,000 documents, and precedent search over a firm’s accumulated knowledge.
What the numbers say Against the base K3 model, Tenet completes almost twice as many held-out tasks on Harvey’s Legal Agent Benchmark (LAB) and 20% more on LAB: Contracts, lifting all-pass rate by 9 and 2 percentage points respectively.
Harvey reports state-of-the-art on LAB: Contracts and second place on LAB, using base-model scores from Vals.The more interesting result is transfer.
Tenet also improves substantially on Mercor’s APEX Agents (corporate law) and Crosby’s Redline Bench — neither seen during training — while holding performance on knowledge benchmarks including LegalBench, CUAD, MAUD, and Scale’s PRBench.Agentic training did not erode textbook legal reasoning.
Cost is co-optimized rather than traded away.Open weights lower price per token; reward shaping that prefers shorter trajectories at equal quality lowers tokens consumed.Harvey reports significant quality gains at stable cost.
How it was trained Training used asynchronous reinforcement learning in sandboxed legal environments built like LAB tasks: a partner-style instruction averaging about 50 words, a client matter of key and peripheral documents, and an expert rubric of atomic pass/fail criteria — roughly 50 per task, hundreds at the extreme.
A single rollout can exceed 1,000 turns.Rollouts are graded by LLM-as-a-judge; ablations settled on Kimi 2.6.Reward combines the fraction of rubric criteria satisfied, a holistic count of legal issues solved, and an all-pass bonus.
The policy is optimized with GSPO using a rank-64 LoRA over the full K3 network, eight task groups of eight rollouts per optimizer step, across ~1,750 environments and >10,000 rollouts per epoch.
Fireworks co-built trainer and rollout deployments at the kernel level, with token-in-token-out and router replay, to keep a large MoE numerically aligned across training and inference.(function(){var f=document.getElementById("mtp-tenet-frame");if(!f)return; window.
addEventListener("message",function(e){var d=e&&e.data;if(!d||typeof d.tenetHeight!=="number")return; if(d.tenetHeight>200){f.style.height=d.tenetHeight+"px";f.height=d.
tenetHeight;}});})(); Three capabilities trained separately Harvey team also post-trained specialist models that Tenet can route to as tools or sub-agents: M&A diligence: On LAB: Diligence, a single task can traverse up to 80M tokens; no baseline passed more than 43.8% of criteria.
With Baseten, Harvey moved to a Recursive Language Model harness where a root agent holds the dataroom in a REPL and delegates to sub-agents.A GLM-5.2 orchestrator alone reached 46.1%; post-training it in that harness via self-distillation reached 60.1%.
Review Table: With Applied Compute, a post-trained GLM-5.2 improved answer quality by 3.6 points and citation quality by 12.1 points at roughly one-tenth the cost per cell, learning to abstain when a question does not apply.Firm knowledge: With Engram, a Qwen3.
8-27B model studies ~100M tokens of client matters into 1M tokens of structured knowledge plus parametric memory.Criteria pass rate rose more than 15%, tokens in completed trajectories fell 58%, and cost per query dropped roughly 90% — 190.8 intelligence-per-token versus 129.
3 for the best frontier configuration.Marktechpost Independent Test Facts #mtp-reality-check-harvey-tenet{background:#111!important;color:#d6d6d6!important;border:1px solid #2a2a2a!important;border-radius:6px!
important;font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Helvetica,Arial,sans-serif!important;line-height:1.55!important;margin:28px 0!important;padding:0!important;overflow:hidden!important;-webkit-font-smoothing:antialiased} #mtp-reality-check-harvey-tenet *{box-sizing:border-box!
important} #mtp-reality-check-harvey-tenet hr,#mtp-reality-check-harvey-tenet p:empty,#mtp-reality-check-harvey-tenet del,#mtp-reality-check-harvey-tenet s{display:none!important} #mtp-reality-check-harvey-tenet a{color:#76B900!important;text-decoration:none!
important;border-bottom:1px solid rgba(118,185,0,.35)!important} #mtp-reality-check-harvey-tenet a:hover{color:#8fd400!important;border-bottom-color:#8fd400!important} #mtp-reality-check-harvey-tenet .rc-bar{background:#161616!important;border-bottom:1px solid #2a2a2a!important;padding:18px 20px!
important;display:flex!important;justify-content:space-between!important;align-items:center!important;gap:14px!important;flex-wrap:wrap!important} #mtp-reality-check-harvey-tenet .rc-eyebrow{color:#76B900!important;font:700 10px/1 ui-monospace,SFMono-Regular,Menlo,monospace!
important;letter-spacing:.18em!important;text-transform:uppercase!important;margin:0 0 7px!important} #mtp-reality-check-harvey-tenet .rc-t{color:#fff!important;font-size:18px!important;font-weight:700!important;margin:0!important;line-height:1.3!important} #mtp-reality-check-harvey-tenet .
rc-src{color:#8a8a8a!important;font-size:11.5px!important;margin:6px 0 0!important} #mtp-reality-check-harvey-tenet .rc-chip{background:rgba(214,69,69,.14)!important;border:1px solid #d64545!important;border-radius:4px!important;padding:9px 14px!important;text-align:center!important;flex:0 0 auto!
important} #mtp-reality-check-harvey-tenet .rc-chip b{display:block!important;color:#e06565!important;font:700 22px/1 ui-monospace,Menlo,monospace!important} #mtp-reality-check-harvey-tenet .rc-chip span{display:block!important;color:#9a8080!important;font:600 9px/1.5 ui-monospace,Menlo,monospace!
important;letter-spacing:.12em!important;text-transform:uppercase!important;margin-top:5px!important} #mtp-reality-check-harvey-tenet .rc-strip{display:flex!important;flex-wrap:wrap!important;gap:0!important;border-bottom:1px solid #2a2a2a!important;background:#131313!
important} #mtp-reality-check-harvey-tenet .rc-cell{flex:1 1 20%!important;padding:13px 10px!important;text-align:center!important;border-right:1px solid #222!important;min-width:88px!important} #mtp-reality-check-harvey-tenet .rc-cell:last-child{border-right:0!
important} #mtp-reality-check-harvey-tenet .rc-cell b{display:block!important;font:700 19px/1 ui-monospace,Me
Related
相關文章
神秘“牛來”大模型上線即登頂 背後廠商至今未揭曉
近日,一款代號為Ox Alpha的匿名AI模型在OpenRouter悄然上線,短時間內調用量迅速攀升,成功衝至平臺榜首,並刷新了該平臺的單日模型用量紀錄。不過,儘管表現驚豔,Ox Alpha背後的開發主體至今仍未揭曉。

AI辦公助手,沒有葵花寶典:五款應用萬字實測報告
AGI-Signal2026.08.24 09:12 · 來自北京全文11936字單項冠軍各有其人。2026年上半年,AI辦公賽道發生了一個根本性變化,工具不再滿足於當“對話框”,而是試圖接管完整任務,寫一段文案、做完一份報告、生成一份PPT,甚至跨應用操作。

Anthropic新模型偷「吃瓜」,最強Fable 5爆冷
Anthropic 近日推出新款 AI 模型,在內部測試中意外展現「吃瓜」能力,引發社群熱議。該模型不僅能快速理解網路迷因與流行語,更在特定任務上表現出人意料,讓原本被外界視為最強對手的 Fable 5 爆冷落後,業界對這項結果感到相當驚訝。目前 Anthropic 官方尚未針對模型實際表現與測試細節做出完整說明,市場則持續關注後續可能的技術更新與應用方向。
Kimi K2.5 月底退役:月之暗面第一代萬億參數多模態模型謝幕
月之暗面官宣第一代萬億參數多模態模型Kimi K2.5將於本月底結束服役。該模型今年1月推出並開源,是Kimi迄今最全能模型,採用原生多模態架構,支持視覺與文本輸入、思考/非思考模式、對話與Agent任務,在Agent、代碼、圖像、視頻及通用智能取得開源SOTA。K3將接力,參數規模再上臺階。
光子躍遷亮相BIRTV 2026:以"AI+影像"重構創作範式,三大板塊解碼下一代影像生態
8月19日,BIRTV 2026(北京國際廣播電影電視展覽會)在北京拉開帷幕。在這場匯聚全球廣電與影像領域頂尖技術與創意的盛會上,光子躍遷以"AI+影像"為核心敘事,攜個人智能影像生態重磅亮相,向行業展示了一個由AI驅動、以人為中心的影像未來。與行業展會常見的深色科技風不同,光子躍遷的展臺以純淨白色為主基調,輔以品牌藍色進行點睛點綴,在千篇一律的深色展臺中脫穎而出,傳遞出品牌年輕、活力、面向未來的基因。

諾亦騰機器人發佈 HiPHI,開源 617.5 小時高精度人體運動數據
作者:潞源 責編:潞源 評論: 8 月 23 日消息,諾亦騰機器人在 2026 世界機器人大會期間發佈 HiPHI,這是一套面向人形機器人學習、數字人,以及計算機圖形學領域研究人員和工程師的高精度光學動作捕捉數據集。據報道,該數據集總長 617.