史丹佛團隊發表 Paper2Agent:將研究論文轉為 AI 代理,重現結果並處理新資料

2026年9月16日 21:59
站內 AI 整理稿

Computational papers ship code that readers must clone, install, configure and debug.That cost keeps useful methods locked inside PDFs.A Stanford team led by Jiacheng Miao and James Zou proposes a fix.Paper2Agent was published in Nature on 16 September 2026.

It converts a paper and its codebase into a Model Context Protocol (MCP) server.Any MCP-compatible agent, such as Claude Code, can then run the paper’s methods through natural language.The authors describe the result as a virtual corresponding author.Is it deployable?Yes.

The code is MIT-licensed and installs as a skill for Claude Code or Codex.Prebuilt AlphaGenome, Scanpy and TISSUE servers run on Hugging Face Spaces.A hosted version is also available at paper2agent.ai.How the Pipeline Works Paper2Agent runs on Claude Code’s agent SDK.

A central orchestrator dispatches specialized sub-agents through 6 steps: Locate and download the codebase.An environment manager builds an isolated virtual environment.A tutorial scanner indexes usable tutorials.A tutorial executor runs them end to end and records reference outputs.

A tool extractor turns tutorials into parameterized MCP tools, and a test verifier validates them.The orchestrator assembles validated tools into 1 MCP server.The validation gate is strict.A tool passes only when expected files appear and numbers match within 3%.

Figures must also match references by perceptual hash, with Hamming distance under 20.The verifier gets up to 6 attempts per function.Tools that keep failing are excluded from the final server.Each server exposes 3 components.

MCP tools wrap the paper’s methods as executable functions: MCP resources hold the manuscript, code links, datasets and figures.MCP prompts encode multi-step workflows, such as the correct Scanpy preprocessing order.The research team used Claude Sonnet 4 for all Paper2Agent applications.

Interactive Explainer #mtp-p2a-embed{max-width:980px!important;margin:24px auto!important;padding:0!important;background:transparent!important;border:0!important} #mtp-p2a-embed iframe{width:100%!important;border:0!important;display:block!important;background:transparent!important;min-height:0!

important} #mtp-p2a-embed p:empty,#mtp-p2a-embed br,#mtp-p2a-embed hr,#mtp-p2a-embed del,#mtp-p2a-embed s{display:none!important} (function(){window.addEventListener("message",function(e){var f=document.getElementById("mtp-p2a-frame");if(!f||!e.data||!e.data.mtpP2A||e.source!==f.

contentWindow)return;var h=parseInt(e.data.h,10);if(h>200&&h<6000){f.style.setProperty("height",h+"px","important");}});})(); AlphaGenome Agent Results For AlphaGenome, Paper2Agent built 22 tools in about 45 minutes for US $14.All 22 passed validation without human intervention.

The team compared the agent with Claude Code plus repository access (Claude + Repo) and Biomni.BenchmarkPaper2AgentClaude + RepoBiomni15 tutorial-derived queries98.7 ± 1.3%82.7 ± 3.4%37.3 ± 4.0%15 novel queries100.0 ± 0.0%78.7 ± 4.4%56.0 ± 3.4%30 open-ended queries82.7 ± 2.4%56.7 ± 2.3%72.2 ± 2.

2% Results span 5 runs, graded by 2 human experts with 96.7% inter-rater agreement.On tutorial queries, median runtime fell 1.9× versus Claude + Repo and 3.1× versus Biomni.The gains persisted when the baseline was upgraded to Claude Opus 4.6.

The agent also re-examined an LDL cholesterol variant, chr1:109274968:G>T.It ranked SORT1 as the likely causal gene.The original AlphaGenome paper emphasized CELSR2 and PSRC1.GTEx shows significant liver eQTLs for all 3 genes.

The research team say this shows how hard causal gene assignment is at such loci.Scanpy, TISSUE and Scale Tests The Scanpy agent received 7 validated tools in about 45 minutes for US $13.On 4 public datasets, it matched human researchers on cell counts, gene counts and top marker genes.

A TISSUE agent reproduced human results on spatial transcriptomics data.Scale tests covered 3 corpora with no manual cleanup: 100 bioRxiv computational biology papers: 74 were agentified, and 593 of 599 proposed tools passed validation.300 questions: Paper2Agent scored 91.2%, versus 80.

3% (Sonnet 4) and 86.3% (Sonnet 4.6) for Claude + Repo.Cost per query: US $0.20 and 1.6 minutes, compared with US $0.38 and 4.3 minutes.10 non-biology papers, including TabPFN, SAM 2 and SAELens: 98.1% accuracy on 42 execution tasks.26 data-focused papers: resource layer 89.0% versus 82.

0% for browser use, 34× cheaper and 15× faster.Paper2Agent also rejected 100% of out-of-scope queries in a permuted benchmark.It recovered from injected dependency, file-path, typo and deprecated API failures.

Paper Agents Collaborating The research team connected 3 agents: AlphaGenome, an MPRA-coupled scCRISPRi screen and a CD4+ T cell Perturb-seq dataset.AlphaGenome flagged GPR137 at psoriasis locus rs887314, with an RNA-seq quantile score of 0.997.

The AI co-scientist proposed 10 validation strategies, and a researcher picked signature correlation.Only GPR137 knockdown matched the CRE perturbation signature.The match appeared under stimulation: Spearman 0.613 at Stim8hr and 0.630 at Stim48hr.

BAD and 3 other candidates showed no significant correlation.A second study paired AlphaGenome with an ADHD GWAS and nominated rs1626703 among 209 candidates.That hypothesis still needs experimental validation.

Key Takeaways Paper2Agent converts papers and repos into tested MCP servers with tools, resources and prompts.The AlphaGenome agent took about 45 minutes, cost US $14, and scored 100% on novel queries.74 of 100 bioRxiv papers were agentified, with 593 of 599 tools validated.

3 paper agents jointly supported GPR137 as the probable psoriasis causal gene.The code is MIT-licensed, with prebuilt MCP servers on Hugging Face Spaces.Check out the Paper and Repo.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data appeared first on MarkTechPost.

Related

相關文章

量子位生成式AI

無問芯穹與華環電子簽署戰略合作,共同探索國產異構算力AI基礎設施新方向

無問芯穹與華環電子簽署戰略合作協議,雙方將結合各自在AI軟體平台、網路通信與硬體研發的優勢,共同探索國產異構算力基礎設施的協同方案。此次合作聚焦於智算中心解決方案及「Token工廠」新模式,目標是推動計算、網路與AI原生基礎設施深度融合,為AI規模化應用提供高效穩定的支撐。

9 分鐘前
IT之家生成式AI

優步全球範圍裁員 10%,被裁員工稱 AI 已大舉滲透日常工作

作者:清源 責編:清源 評論: 9 月 18 日消息,據《商業內幕》今天(18 日)晚間報道,在優步(Uber),AI 已經滲透到員工工作的許多環節,從回答 Slack 裡的內部問題,到替乘客行程中聯繫客服時收到的消息撰寫回復。6 名近期遭裁員的員工透露,過去幾個月,AI 在工作中的使用範圍明顯擴大,其中一些人甚至會通過提示詞讓 AI 完成相當一部分任務。

3 小時前
鈦媒體生成式AI

月之暗面遞表之後,Kimi 的成色要被驗算三遍

舒澤品牌手記2026.09.18 18:16 · 來自浙江全文4982字00:00 / 14:05Anthropic 的 30 萬次指控,會成為招股書的第幾頁?文 | 舒澤品牌手記9月17日,月之暗面發佈了一套金融行業解決方案。按官方披露,中信建投、中金公司、易方達等數十家金融機構已經在用 Kimi 處理投研建模、風險排查和盡調材料——研究人員把管理層報表、審計報告和盡調文件交給 Kimi,拿回一份可以繼續調整假設的 Excel 模型。同一天,深圳商報記者就港股上市進展、股東架構調整等事項向月之暗面發去採訪函。

5 小時前

Calibre上手 AI 互動寫作:電子書管理器搖身變成"文字冒險遊戲引擎"

這個遊戲默認藏而不發,不會跟著 Calibre 啟動就冒出來。用戶得主動在"首選項 — 工具欄和菜單"裡把它請到主工具欄,才算真正激活。它的玩法很清晰:由 AI 在後臺搭起並掌管一個虛構世界,用戶通過不斷輸入文字來推著故事往前走,等於把"讀電子書"這件事,翻轉成了"和 AI 一起寫故事"。

7 小時前