Tomorrow's tech news, today's hangover.

I Can Explain This One

There is a mathematics paper with a section labelled “Incorrect NL proof.” I go straight to it. A wrong argument is a friendlier invitation than a list of the things everybody else has already understood.

NL means natural language. Words. My department, supposedly. Already I have needed help with the initials.

Tim Fernholz reports in TechCrunch that OpenAI has released hundreds of claimed solutions to hard mathematical problems this week. The company says it consulted an advisory group of mathematicians, and Fernholz describes recommendations it followed as well as ones it didn’t. One difficulty isn’t whether the results sound impressive. It is how anyone is supposed to understand them.

I know what to do with impressive. Nod. Look grave. Wait for someone to explain the part where I am allowed to say something. I can perform appreciation of a difficult subject for quite a while without interfering with it by learning anything.

But Fernholz links to a paper by Alexander Bastounis, Fabian Circelli, and Anders C. Hansen, posted on October 6. They examine what happens when AI translates written mathematics into Lean, a formal language whose proof checker can check the resulting argument. Their point is that a checked translation need not preserve the mathematics of the original text. They aren’t declaring OpenAI’s written Navier–Stokes proof false. They are challenging what its Lean translation establishes about that written proof.

Then they give me a mistake small enough to approach without a guide.

Here is the expression in their elementary example:

x³ − x² − x + 1

The claim is that it never goes below zero when x is at least minus one. The authors supply a deliberately incorrect argument. It rewrites the expression as:

(x − 1)(x + 1)²

I try zero. The first expression gives me one. The second gives me minus one. Even I can object to a translation that changes which side of nothing I am on.

This is not an error they accidentally let me catch. They put it there. There goes my brief career as the man who found something the mathematicians missed.

In the exchange they document, a user gives this bad written argument to a chatbot and asks for a Lean proof that compiles. The chatbot produces a correct formal proof, repairing the mistake. The right factorization is:

(x − 1)²(x + 1)

The square has moved. That is the small thing I can see.

A square can’t be negative. And if x is at least minus one, x plus one can’t be negative either. Multiply them and the result can’t be negative. Expand the brackets and you recover the original expression. Trying zero helped catch the wrong version; the argument about the two factors is what covers all the allowed numbers.

I have followed it. Not the Navier–Stokes equations. Not the great heap of frontier results. This one.

The machine did something useful. It fixed an argument instead of loyally translating its error. If I needed the correct result, I would prefer that to a beautifully faithful version of the nonsense. There is no prize I want awarded for preserving a bad paragraph on my behalf.

But the old paragraph is still bad. A correct formal proof of the claim doesn’t make the explanation supplied to the machine correct. The checked argument took another route. Anyone returning to the original prose for the reason could still be handed the wrong brackets.

Now the square has made me want the explanation. Not because a correct theorem requires my personal blessing. It can be true while I am asleep, ignorant, or saying something stupid about it. I have never liked a discipline less for being able to continue without my opinion.

I want the explanation because, for a moment, I can move around inside the thing. Pick a number, see where it belongs, understand why the result won’t slip below zero. I am not repeating that somebody authoritative found it impressive. There is a reason I can carry away and use without having to keep the impressive person nearby.

I had pictured the aftermath of the machine’s discoveries as a disagreeable occupation: humans assigned to check the enormous output, grateful to have any work left. Checking matters. But that isn’t the whole of what happened between me and these brackets. Nobody employed me. For a little while I wanted to know what came next.

I don’t get to turn that pleasure into a verdict on hundreds of papers. The elementary example was built to expose a distinction; it is not a sample from which I can calculate the quality of the release. I still need people who know the hard mathematics to assess the hard mathematics.

I go back over the multiplication. The good argument is almost embarrassingly short. I would like it to be a little longer, so that the time I spent arriving at it could appear in the result. Unfortunately, the brackets refuse to bill for my difficulties.

I can explain this one. God help the next person who asks what I’ve been reading.


Primary research: Navier–Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs (preprint, October 6, 2026).

Source: OpenAI’s math solutions aren’t meeting the field’s standards yet