Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Back to Articles Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents Team Article Published September 29, 2026 Upvote 2 Antonio Tiene AntonioTN Follow MultiverseComputingCAI Ander Alvarez Sanz ander-alvarez Follow MultiverseComputingCAI Oliver Wirjadi oliverwirjadi Follow MultiverseComputingCAI Tool-using LLM agents no longer read from a single retrieved passage.
Through the Model Context Protocol (MCP), an agent can call a search tool, inspect a structured patient or account record, query a database, and pull metadata, then weave all of it into one answer.That m aakes the usual question of factuality more subtle than it looks.
Most of the systems built to check LLM answers, from RAGAS faithfulness to fine-grained checkers like MiniCheck, AlignScore, and SummaC, ask whether a claim is supported by the available evidence once that evidence has been pooled together.
In their usual form, they do not tell us which MCP tool output supports each claim, or whether that is the source the answer names.
Our latest paper, ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (read it on Hugging Face, or on arXiv in the meantime), targets that gap.
The failure mode we care about is one we call cross-source conflation: a claim that is true somewhere in the evidence, but attributed to the wrong source.A source-blind verifier may pass it, because the fact does exist in the pool.A source-aware verifier should not.
The problem: supported somewhere is not the same as supported by the right source Consider a customer support agent that answers, "According to the account record, this plan includes a 30-day refund window.
" The refund window may be perfectly real, but stated in a policy document, not in the account record the answer points to.Pool the two together and the claim looks supported.
Keep them separate and the attribution is wrong, and in a data-sensitive setting a wrong attribution can be as damaging as a wrong fact.
The same pattern shows up in a clinical agent, where a patient-specific medication detail taken from a patient-history tool becomes misleading the moment the answer presents it as a finding from the medical literature.
A claim can be supported by one MCP source while the answer attributes it to another.Source-blind scoring sees support in the pooled evidence and passes it; ProvenanceGuard separately checks whether the supporting source matches the one the answer states or implies.Source: paper Figure 1.
This is why faithfulness scores, useful as they are, are not enough for MCP agents.An answer carries provenance, sometimes explicitly ("according to the account record") and sometimes implicitly.ProvenanceGuard keeps that connection between claim and source available for inspection.
What ProvenanceGuard does ProvenanceGuard is a post-generation verification layer that sits on top of a black-box MCP agent.It runs after an agent produces an answer, and never collapses the evidence into one anonymous context.Instead it carries the source identity all the way through the pipeline.
It reads the captured MCP trace, including the tool outputs and their source IDs, without retraining the agent.
Then it does five things in sequence: it breaks the answer into specific claims, finds the source most relevant to each one, checks whether that source actually supports it, compares the source with the one the answer names or implies, and finally emits both a per-claim source verdict and a global, answer-level allow or block decision.
The verification flow.Source identity is preserved through decomposition, routing, support scoring, attribution checking, and repair, rather than being pooled.Blocked answers can go through RARR-style repair and be re-verified.Source: paper Figure 2.A few of the design choices are worth calling out.
For the experiments in our paper, we used local models so the captured traces could be processed in a controlled, offline setup: MiniLM helps find the relevant source, a DeBERTa NLI verifier model checks whether that source supports the claim, and a local language model helps break answers into claims.
The verifier also checks literal values closely: a number, date, or identifier absent from the source cannot pass merely because the sentence sounds plausible.A calibrated decision step combines these signals.
If an answer is blocked, a RARR-style repair step can try a source-grounded revision or a safe fallback, which the verifier then checks again.Those named models are the setup we evaluated, not a requirement of ProvenanceGuard.
The same claim, source, and decision steps can be adapted to hosted models where a team prefers cloud services; a new setup would need its own testing and calibration.Our reported results come from the local configuration.
Its conservative decision policy suits data-sensitive review, where getting the source right matters more than producing the fastest possible answer.Results We tested ProvenanceGuard on answers from a medical agent that had used patient records, research articles, and other tools.
This gave us 281 real traces to study.Medicine is a useful test because a fact from a patient's record and a fact from general research cannot be treated as the same source.The method can also be used in other fields when an agent keeps a record of its tool outputs and source IDs.
For the main test, human experts checked 361 claims from 40 answers set aside from the data used to develop the system.The most direct result is this: experts said 139 claims should not pass, and ProvenanceGuard caught 138 of them.It let one through.
It also held 67 claims that the experts considered supported, sending them for review or repair.This reflects the cautious setting we tested: it favors a second look at some supported claims over letting unsupported ones through.
For claims with an identifiable source, it also picked the right source about 86% of the time in this test.We ran four other support checkers on the same claims.
ProvenanceGuard scored highest on the paper's measure of how well a system catches claims that should be blocked while avoiding unnecessary blocks.The other checkers in this comparison did not tell us which tool output supported each claim.
ProvenanceGuard records that connection, so a reviewer can see the source checked for each claim and the decision it produced.Verifier Reject/block F1 Emits claim-to-source ID ProvenanceGuard (ours) 0.802 Yes MiniCheck 0.783 No RAGAS Faithfulness 0.758 No AlignScore 0.662 No SummaC-ZS 0.
436 No Binary support metrics on the same held-out claim packet.ProvenanceGuard matches or beats the source-blind baselines on blocking while also producing per-claim source verdicts.Source: paper abstract and Table III.
Checking claims when sources look similar In a separate, harder test with several similar sources, ProvenanceGuard scored 0.846 F1 for deciding which claims to block, but identified the exact source correctly in 50.3% of claims.Telling similar sources apart remains an important area for improvement.
We also ran a controlled test focused on wrong attribution: we changed the named source in 50 cases while leaving the supporting evidence intact.ProvenanceGuard caught all 50 swaps.
This shows it can detect a clear source error, while the harder test shows the challenge of choosing among many plausible sources.Repairing blocked answers Blocking is only useful if there is something to do with a blocked answer.
Wired to the RARR-style repair loop, the full-trace run resolved all 173 blocked answers, though 144 of them ended in fallback text rather than a substantive rewrite, which is the system choosing to avoid an unverifiable answer rather than manufacture one.
On reconstructed multi-source test traces, a fresh repair run resolved all 59 initially blocked answers with only two terminal fallbacks.
As an offline gate the overhead is modest, roughly half a second per answer on the reported local configuration, with the NLI and routing calls themselves in the tens of milliseconds.
Why this fits Multiverse Computing As agents move from single-passage RAG to multi-tool MCP setups, the question of which source a fact actually came from stops being a footnote and becomes part of what factuality means.ProvenanceGuard makes that source connection visible claim by claim.
For Multiverse Computing, that means a way to check existing agents while keeping sensitive traces in a controlled environment when needed.The medical study is one use case; the same approach can be adapted wherever an agent's trace preserves its tools and sources.
That adaptation is already visible in NVIDIA NVFlow, which merged an optional grounding-verification stage for its finance agent.It checks completed answers against the SEC excerpts the agent retrieved and saves separate decisions without changing the original rollout or training data.
The NVFlow contribution uses ProvenanceGuard's source-aware verification approach; the repair loop discussed above belongs to the broader research system.ProvenanceGuard was also presented as a poster at the Agentic AI Summit 2026 at UC Berkeley.
Want the full technical details, including the routing and NLI derivations, the calibration ablations, the multi-source stress slices, and the complete results tables?
Read the full paper on Hugging Face, or get in touch with our team to talk about applying source-aware verification to your own agents.
Models mentioned in this article 2 Papers mentioned in this article 1 More from this author Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem 33 September 21, 2026 Safety for Whom?
Refusing the Right Subset of a Topic, Not the Whole Topic 31 September 8, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 2 Models mentioned in this article 2 Papers mentioned in this article 1
Related
相關文章

奔赴南海科考,我國自主研發 4500 米級深海作業機器人啟航
作者:浩渺 責編:浩渺 評論: 感謝網友 很宅很怕生 的線索投遞!9 月 29 日消息,據央視新聞報道,9 月 28 日,在山東威海港,我國自主研發的 150 馬力 / 250 馬力 4500 米級深海作業機器人正式啟航,聯合哈爾濱工業大學(威海)奔赴南海開展聯合科考。

“AI空頭”思考:雲大廠“ROIC下滑”、大模型“難有護城河”、OpenAI是“AI時代WeWork”、Muse背後“人肉電池”
兩位“AI空頭”認為,超大規模雲廠商增量ROIC已於2024年見頂並快速下滑,若趨勢延續,2027年中期將跌破資本成本。近日,知名做空機構Chanos & Co創始人Jim Chanos與紐約大學榮譽教授、AI研究者Gary Marcus近日在RiskReversal播客節目中展開深度對話。

ARR真的是大模型企業的黃金指標嗎?
數智奔流2026.09.29 19:13 · 來自河南全文4129字00:00 / 13:08快速上漲的ARR,要花多少錢才能維持。文 | 數智奔流,作者 | 熵野,編輯 | 沈校9月,幾件看似沒有關聯的事件,揭露了大模型公司正在發生的變化。

82億美元,硅谷上演頂級版“Girls help girls”
山農下山2026.09.29 18:12 · 來自甘肅全文4164字蘇姿豐與李飛飛“會師”。文 | 山農下山溫柔的收購與現實的考量“去尋找一個更新的世界,永遠不會太晚。”AMD在9月28日官宣收購李飛飛創辦的 World Labs,82億美元、全股票交易。

安卓端Google Assistant正式謝幕,Gemini接管手機語音入口,手錶汽車卻還留著舊管家
9月29日據外媒報道,谷歌從上月起就開始向用戶群發郵件,預告手機端的 Google Assistant 即將停用;如今大批安卓用戶反饋,他們的手機已經實實在在調不動那個熟悉的語音助手了。一些用戶更發現,這幾天 Gemini 應用的賬戶菜單裡,切換到 Google Assistant 的入口已經消失——從選項到服務,舊助手正在被一點點抹掉。

Claude Sonnet 5.5發佈:性能逼近Opus 5.5,價格減半,安全防護升級
更快、更便宜了。在推出旗艦級模型Claude Opus 5.5僅一週後,Anthropic又帶來了Claude 5.5家族的第二款產品。當地時間9月28日,Anthropic正式發佈Claude Sonnet 5.5。作為面向日常辦公、軟件開發等場景的主力模型,Sonnet 5.