The AI Architect

The AI Architect

OpenAI Solved 10 Open Math Problems for $2,000. Two Mathematicians Recognised Their Own Work in the Proofs

Two thousand dollars, 249 pages, every logical step certified by machine. And two mathematicians who recognise their own uncited arguments. The check ran. It was not looking there

Matija Vidmar's avatar
Matija Vidmar
Aug 09, 2026
∙ Paid

AI Architect Weekly — August 3–9, 2026

Three levels, one question: who answers for what AI produces when the automated check says everything is fine.

This Week

Two thousand dollars in tokens and 249 pages of formal proof: that’s the tab OpenAI published in early August, alongside ten open math problems solved by a model nobody outside the company can run. In the same seven days, Palantir reported 93% growth and generated $2.1 billion in cash selling something that looks nothing like a language model. And in Washington, five public manifestos and one framework that will stay classified tried to put in writing which object the government should control before it reaches the market.

Three stories that look like they belong to different worlds, a research lab, a public company, a Senate committee room, and instead share the same object seen through different lenses: an AI output, a commercial contract built on top of that output, a legal definition trying to frame it before it becomes a product.

The thread isn’t obvious from this week’s headlines. It gets clearer the moment you ask, story by story, who actually put their name on the line.


THE CHECK

Lean verified every single logical step in the ten proofs. It could not verify, because it can’t, whether the arguments had already been written by someone else.

In Lean, the formal proof language mathematicians use to have a machine check every step of an argument, there’s a word that marks surrender: you write “sorry,” and it means that step wasn’t proven, it was just assumed. In the repository OpenAI uploaded to GitHub under an Apache 2.0 license, inside a 249-page manuscript and the Lean 4 proof certificates for all ten proofs, the sorry count is zero. Zero across ten problems each open for at least a decade, zero steps left hanging, everything checkable by anyone who downloads the code and runs it. Then two mathematicians opened the PDF, and recognised their own arguments. The counter read zero because it was counting one thing only: whether step N follows from step N-1. It wasn’t counting whether that argument had already been written, by someone else, years earlier.

The second jolt came with the bill: roughly $2,000 in tokens to generate all ten published solutions, the compute stack of a company worth hundreds of billions reduced, for once, to a figure that fits on an expense report. Behind that number sits a scope that widens the closer you look: ten open problems in mathematics and theoretical computer science, each stuck for at least a decade; the explicit construction of a non-sofic group, a question open for 27 years since Mikhail Gromov introduced the notion of sofic groups in 1999; a counterexample to Connes’ rigidity conjecture; improved bounds on sphere-packing density in high dimensions; several Erdős problems. OpenAI also cites high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice-based cryptography and extremal combinatorics among the fields touched by the work.

What Lean certifies is precise, and shouldn’t be confused with something bigger: it guarantees the proof follows logically from the statement it was given. It doesn’t guarantee that formal statement is a faithful rendering of the problem mathematicians believed was open, and it doesn’t guarantee the argument is original. Those are two different questions, and only the first has an automated checker.

Steven Miller, a mathematician at Yeshiva University, says the sphere-packing proof reuses an argument from his 2016 paper without citing it. Fournier-Facio, at Cambridge, flags the same pattern, unattributed prior work, in the non-sofic group result. These are claims, not verdicts, and should be read that way. But neither of them found it by running a checker. They found it by reading. There’s a second, quieter limit too: the papers were written by human researchers who used Lean to formally verify arguments generated by the model, but nobody outside OpenAI can run that model. The result is checkable line by line. It isn’t reproducible by another lab, a gap I wrote about when covering the hidden cost of amateur AI: the bill always arrives, even when it’s certified line by line.

This matters beyond anyone proving theorems. Every AI output today sails through the checks that can be automated, the code compiles, the numbers add up, the text is coherent, the cited source actually exists, and quietly fails the checks a machine can’t run: is it original, is it the right question, has someone already done this without saying so, whose name is on it if it’s wrong. Astra just staged that gap at the largest scale possible, with the best verification apparatus in the world, and the gap stayed.

Astra, OpenAI’s next major model family, announced August 2-3, isn’t released. The version that solved the ten problems is an internal prototype. Sam Altman has already shown it to policymakers in Washington, and Astra is expected to be the first test case of the Trump administration’s pre-review framework. A model nobody outside the company can run is about to become the test of an oversight system built for models that, by definition, have already shipped.

The sorry count read zero. Not because there was nothing left to check, but because it was checking something else entirely.

Lean can tell you the proof is correct. Nobody has written the program that tells you it is yours.

Take the last piece of AI-assisted work you shipped and write down which checks you actually ran. Almost always there are three: the numbers add up, it reads coherently, it’s in the right format. Now add the fourth question no tool will ever ask for you: where did it come from? If you can’t answer that, you haven’t verified the work. You’ve verified its shape.


Also This Week

Jeff Dean leaves Google after 27 years The chief scientist is departing, taking Sanjay Ghemawat, Oriol Vinyals and Quoc Le with him to found Discovery Loop, which will use AI to automate scientific and engineering processes. Alphabet lost about 4% on the news.

AI search is sending Shopify more traffic, not less Shopify says AI search is boosting both traffic and sales, the opposite of what’s happening to online publishing.

Nvidia weighs cutting memory on Rubin Ultra Internal prototypes carry 192-256GB of HBM versus the 1TB HBM4E configuration shown in 2025, blamed on HBM shortages; launch expected late 2027.

An AI designed 16 viable viruses that never existed in nature Stanford and Arc Institute’s Evo model generated 700,000 genomes, synthesized 285 as DNA and produced 16 viable viruses able to infect bacteria, in a study published in Science on August 6.

An economist argues AI is too expensive to replace workers Steve Hanke: running advanced AI systems costs more than employing people once you count compute, electricity, water and infrastructure.


What you just read is a third of the week, and it’s the third about a PDF. The other two stories are the same verification gap where money and rules actually move: a company that generated $2.1 billion in cash in one half-year selling the check no machine knows how to run, and a government writing rules around the one object that did nothing this week. One question stays open here, who signs off when AI gets it wrong inside your company. Subscribe and read the other two: fifteen minutes, issue seven.

User's avatar

Continue reading this post for free, courtesy of Matija Vidmar.

Or purchase a paid subscription.
© 2026 Matija Vidmar · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture