ATRIUMsearch → argument graph
Article · 2026-08-01 · 6 moments

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progre ✦ AI generated

01
Prediction

AI is the catalyst for a fundamental shift in mathematics toward 'big mathematics': large-scale, decentralized collaborations between humans and machines, where complex tasks are diced and sliced, humans claim the creative parts, and AI does the lion's share of the technical grunt work.

Terence Tao, neither dismissive of AI nor fearful of it, sees it as the catalyst for 'big mathematics' — a future of large-scale, decentralized human-machine collaboration in which humans take the creative parts and AI does most of the technical grunt work.

transcript

Article author (via Hacker News): Unlike some of his peers, Tao is neither dismissive of AI nor fearful. Instead, he sees it as the catalyst for a fundamental shift in the discipline—a transition toward what he calls "big mathematics." He envisions a future of large-scale, decentralized collaborations between humans and machines, where complex mathematical tasks can be diced and sliced, with humans claiming the creative parts and AI doing the lion's share of the technical grunt work.

02
Context

Anthropic discovered cryptographic weaknesses using Claude with Mythos Preview, spending $100,000 on tokens with prompts that demanded 'proper research to find genuinly hard findings.'

Setting the scene: days before the OpenAI story, Anthropic had flexed by using Claude with Mythos Preview to discover cryptographic weaknesses, burning $100,000 in tokens on prompts that explicitly demanded genuinely hard research rather than low-hanging fruit.

transcript

Article author (via Hacker News): A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings."

extends · 1

04
Claim

OpenAI set an internal version of Astra, its next major model, on finding solutions to ten mathematical problems that have seen no progress on the main result for at least a decade, spending less than $2,000 at GPT-5.6 Sol token prices on each one.

The main news: OpenAI reports its next major model, an internal version of Astra, solved ten mathematical problems that had seen no progress for at least a decade, each at a claimed cost of under $2,000 at GPT-5.6 Sol token prices.

transcript

Article author (via Hacker News): Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one.

extends · 3

05
Fact

OpenAI's release included Lean 4 formalizations of the results in the openai/ten-proofs repository, a paper describing the solutions, and an LLM-generated PDF reconstructing how each proof came together from the unpublished reasoning traces — a decent level of transparency, though the prompts themselves were not published.

On transparency: the openai/ten-proofs repository holds Lean 4 formalizations, backed by a paper and an LLM-generated PDF that reconstructs the proof process from reasoning traces — decent disclosure, though the author still wants the prompts themselves released.

transcript

Article author (via Hacker News): The openai/ten-proofs repository has Lean 4 formalizations of their results, and there's also a paper describing the solutions and an additional LLM-generated PDF where the model "reconstructs how the proof came together" based on the unpublished reasoning traces. That's a decent level of transparency, but I want to see the prompts they used!

Highlight slides
Related episodes