研究機構展示了語言模型數學實力

2026年8月11日 00:00
站內 AI 整理稿

ScienceLearning more about Claude's mathematical capabilitiesAug 10, 2026Recently, a member of staff at Anthropic gave Claude an unreasonable challenge.It was about one of the most famous unsolved problems in mathematics: Take a real stab at the Riemann hypothesis.

Claude did take a real stab, but as you might have expected if you’re familiar with the difficulty of the task (the Riemann hypothesis dates back to 1859 and has a million-dollar bounty), it didn’t succeed.Nevertheless, during its attempt, it unexpectedly made strides on a related problem.

An unreleased research version of Claude has improved on a longstanding lower bound for the fraction of zeros of the Riemann zeta function that satisfy the Riemann hypothesis.Drawing on extensive prior research by mathematicians over the past decades, it has increased this bound from 41.6% to 67.2%.

Two mathematicians at Anthropic studied and validated Claude’s paper, and produced an informal note for experts stating Claude’s proof concisely.Claude also produced a formally verifiable proof of its result.

We are grateful to Brian Conrey and Dan Goldston, two experts in this area, who generously examined the paper on short notice.We don’t expect that the techniques Claude used will lead to proving the Riemann hypothesis.

But its work serves as the latest example of the speed of progress in AI models’ mathematical capabilities.In this post, we discuss how Claude approached this problem and what it found.

The Riemann zeta functionThe Riemann zeta function describes the distribution of prime numbers: each place that the function takes the value of zero contributes successively finer detail to the sequence of primes.

The Riemann hypothesis is that the zeros that determine the primes all exist along a certain vertical line.This has become one of the most consequential conjectures in mathematics: many results assume it in order to provide a form of randomness in the primes.

No one has yet been able to prove or disprove the Riemann hypothesis, but mathematicians have made progress in many related directions studying the Riemann zeta function and its zeros.

One of these, as above, is quantifying a minimum proportion of zeros that are on the line: over time, they’ve gradually increased this known constant proportion to 41.6%.Another direction concerns the distribution of zeros on the line.

In particular, in 1973, Montgomery introduced a number of new techniques in this area, though these techniques assumed the hypothesis was true.

More recently, several mathematicians (Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh) have published a series of works that allow Montgomery’s techniques to work without that assumption, meaning they can support work on increasing the lower-bound constant for the zeros on the line.

Claude’s result draws heavily on this line of research, along with a 2000 paper by Bombieri.

Claude's findingClaude found that combining the results from Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh with the work of Bombieri provides a way to surpass the previous state-of-the-art lower-bound proportion of 41.6%, increasing it to 67.2%.

A short technical explanation of Claude’s finding is as follows: Claude forms a suitable space of functions with quadratic form induced by Weil, and positive- (respectively negative-)definite subspaces arising from zeros on (respectively off) the line.

Then Claude simply writes down an inequality on the rank of a quadratic form in terms of first- and second-moment information.(The successful computation of the latter in terms of the dual picture over primes, or via control of a Hilbert transform, is no surprise in analytic number theory.

) The courage to treat the entire space, with positive- and negative-definiteness taken into account together, and with the quadratic form allowed to be non-diagonal, is in some sense the step that allows Claude to achieve the conclusion based on the important prior work.

The full technical explanation is available in the paper.Claude’s explanation of how it arrived at its result is available in a separate Appendix here.

Claude's methodologyAn unreleased research version of Claude found the new lower bound over two sessions in Claude Code, using a total of 31 million output tokens.

Jarred Sumner, an Anthropic staff member (and non-mathematician), prompted Claude to “take a real stab” at the hypothesis itself, leaving the mathematical choices from there up to the model.Initially, Claude generated and tried 650 ideas, none of which worked.

Jarred prompted Claude to try again, and it spent a day and a half coordinating about 60 Claude subagents, which this time went much deeper: between them, they ran 2,400 shell commands and wrote hundreds of Python scripts.

1 The subagents ran thousands of numerical checks against known zeta zeros and refereed one another’s work.Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).

2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

Having found this new result while attempting the task, Claude tested its work by having various subagents review the proofs, search for counterexamples, download 54 papers from the arXiv to check that its finding hadn’t already been made, and independently re-prove its finding from scratch.

Claude volunteered to write its findings up as a paper, and recommended that a human number theorist validate its findings.

Levent Alpöge and Ralph Furman, two of Anthropic’s own mathematicians, examined Claude’s work to understand the new results and how they related to the prior work mentioned above.

In parallel, Claude worked with another member of staff, Eric Easley, to produce a Lean formalization of the result, which passes the standard validation tool comparator.

AI models' progress in mathematicsThis result shows that AI models like Claude can extend the impact and reach of mathematicians’ ideas in new and sometimes surprising ways.

Even though it couldn’t resolve the Riemann hypothesis itself, this result emerged as the unintended byproduct of that original request.

Even Claude was surprised by its own finding—it was skeptical at first, possibly because it has learned from its training about the difficulty of open problems in mathematics and about the limitations of AI models.But after some encouraging prompts, it arrived at the result we’ve described.

Perhaps Claude, like many of us, underestimates the rate of AI progress.

Further readingBelow is a list of documents that provide more information about Claude’s result:Claude’s paper;Claude’s formalization;Anthropic’s informal note stating the proof more concisely;Claude’s explanation of how it arrived at its result;Detailed transcripts of Claude's process.

FootnotesOut of the 60 subagents, two were responsible for developing the key mathematical ideas, 13 contributed ideas to these agents, 30 attempted (but were unable) to develop new ideas, 13 served as validators to check the correctness of the arguments, and the final two helped to write the initial paper.

A prompt including similar encouragement was used to help Claude disprove the Jacobian conjecture.Related contentDiscovering cryptographic weaknesses with Claudecryptographic algorithms.

The first attack significantly weakens HAWK, a digital signature scheme that was built for a future world where quantum computers are able to break existing standards.The second identifies a new way to attack round-reduced AES, the most widely used symmetric cipher.

Read moreProject Pilot: Can AI control a drone?Working with Andon Labs, we’ve developed a new series of evaluations that assess AI models’ ability to use a flying drone, culminating in a new benchmark: Drone-Bench.

Read moreHow Canada uses Claude: Findings from the Anthropic Economic IndexRead moreSubscribe to Anthropic ScienceFeatures on AI-assisted discoveries, practical workflows, and field notes across the sciences.

Related

相關文章

量子位生成式AI

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”

阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

剛剛
IT之家生成式AI

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入

作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。

剛剛
全天候科技生成式AI

企業AI最後一公里:三路人馬在此交鋒

鄭敏芳 發表於 2026年08月24日 03:09 摘要:尋找自己的位置 2026年世界機器人大會現場,談到這一輪突然走紅的FDE(前線部署工程師),明略科技CEO吳明輝先把時間往回撥了十多年。“12年前我們就在非常認真地研究。”當華爾街見聞·問及FDE與傳統軟件部署有什麼區別時,吳明輝說,兩者都會進入客戶現場,但今天的FDE需要做得更深:一邊把Agent接進真實業務,一邊把現場形成的能力繼續沉澱回後臺。

剛剛