Back to Explore
NewsVibeFix 编辑部Updated Oct 11, 2026

Terence Tao Declares "War" on OpenAI: 700+ AI-Generated Math Proofs Split the Math World

Fields Medalist Terence Tao led the Association for Human Mathematicians in a joint statement urging mathematicians worldwide to stop collaborating with OpenAI, protesting its one-shot release of 700-plus AI-generated math proofs. Turing laureate Yann LeCun calls it the dawn of a new era for mathematics. A debate over "readability vs. problem-solving power" is splitting the math world — and it is the lived question of everyone in the vibe coding era.

Cover image: Terence Tao vs OpenAI in the AI-generated math proofs debate

The ugliest public breakup in the history of mathematics happened inside an AI company's GitHub repository. On October 9, QbitAI reported that Fields Medalist Terence Tao had led the Association for Human Mathematicians (AHM) in publishing a joint statement urging mathematicians worldwide to stop all collaboration with OpenAI. The statement's fire was aimed at one thing that happened in late September: OpenAI dumped more than 700 machine-generated mathematical proof files into the public at once. Tao's side minced no words — as paraphrased in the reporting: "publishing 700-plus files in one go isn't research at all; it's pure flexing of compute hegemony."

First, the weight of this. Mathematics has never lacked quarrels — priority disputes, school rivalries, fights over proofs — but a joint statement calling on an entire profession to boycott a company is a first. And the target isn't some pseudoscience outfit; it's the strongest AI company alive. "AI for Science" has been chanting "machines will reshape research" for years; this is the first time a top-tier human scholarly community has stood up and said: no, we don't accept your way of doing it. This deserves to be taken seriously — not because of Tao's name, but because the core of the dispute — does something AI generates still count as a "humanly intelligible achievement"? — is the lived question of everyone in the vibe coding era. Let's get the chain of facts straight first.

AHM 抵制 OpenAI 事件时间线(示意图)

1. The trigger: ~8,000 problems, ~5% hit rate, ~3 hours of compute

Piecing together the timeline from QbitAI's October 9 report (verified by our scout on the QbitAI homepage, timestamped "the day before yesterday 08:35"):

  1. Late September: the dump. OpenAI published a batch of mathematical proof files generated by its internal models — 700-plus in total. Three were later withdrawn from its GitHub repo, leaving roughly 719. The sheer scale was itself a declaration: this wasn't a submission, wasn't a paper, it was a dump. The traditional rhythm of math publishing is "one paper polished over years"; OpenAI's rhythm is "700 papers in one release." Two research ethics collided on day one.
  2. Procedural injustice: ignoring a red line it helped draw. OpenAI had previously co-founded an advisory group on mathematics and AI (AGMAI) with the Princeton Institute for Advanced Study, which drew an explicit red line: frontier AI companies should not unilaterally test advanced mathematical problems on their internal models. According to the reporting, OpenAI ignored that warning, ran roughly 8,000 open mathematical problems through its internal models, produced 700-plus proof files at a hit rate of about 5%, and published them. To put it in terms any reader will grasp instantly: it tore up a rule it had helped write. That is the biggest source of the "distrust" in AHM's statement — the issue isn't just that AI solved the problems, but that the solving bypassed the procedures the community had agreed on. Act first, apologize never — and the rule shredded was one it had signed.
  3. October 8–9: the AHM joint statement. Led by Tao, the statement urges mathematicians worldwide to stop collaborating with OpenAI. Cross-reporting from CNA/CNYES and the DEV community (both published two days ago) confirms the story's propagation chain — this isn't one Chinese outlet's exclusive spin; it's a brewing international controversy.

But what really keeps mathematicians up at night is the cost ledger. Laid out as a table, the numbers hit harder:

DimensionHuman scholar (traditional path)OpenAI's internal model (as reported)
Time to crack a hard problemYears — sometimes a lifetime of research~3 hours of GPT-Pro-class compute
Output per releaseOne paper, polished over years~8,000 problems brute-forced into 700+ proofs, published at once
Form of outputHuman-readable, peer-reviewable, teachable proofMachine-generated files, no peer review

A problem a scholar might chew on for years, even a lifetime, gets "solved" by a machine in an afternoon — and it hands in 700-plus of them in one go. This is no longer competition; it's a victory party thrown on your head after a dimensionality-reduction strike. At least half the fury in Tao's "flexing compute hegemony" line comes from here: when the cost of solving approaches zero, the academic dignity built over generations loses its pricing power overnight. What you spent your life proving, someone else "proved" with a few hours of electricity — anyone would have to ask: what exactly have I been doing all my life?

2. Point vs. counterpoint: Tao vs. LeCun

The most valuable part of this debate is that the two sides aren't even arguing on the same dimension. Hear out Tao's camp first.

Tao's camp: a "proof" nobody can read — is that a proof at all?

Quality is the second pillar of AHM's statement. Theoretical computer scientist Scott Aaronson called the scene "Mathocalypse" — as paraphrased in the reporting, the word doesn't mean "AI is too strong, humans are finished"; it means human mathematics, as an intelligible and transmissible body of knowledge, is facing its apocalypse. His wife, unique-games-conjecture researcher Dana Moshkovitz, was blunter — as paraphrased in the reporting, she said the proofs read "like they were written by someone on hallucinogens": leaps of logic, piled-up prior conclusions, bizarre constructions no human has ever seen. These files went through no peer review and offer no readability — they may be formally correct, but no human can actually "read" them.

Tao's own warning goes further: "proof indigestion" and irreversible academic pollution. His logic deserves to be chewed on word by word — mathematical research was never just about "getting the answer"; the path of exploration is itself knowledge. Once an open problem is declared "solved" by a machine, it becomes pointless for humans to research along the original path: you haven't even set out, and someone else's flag is already planted at the destination. Worse, the pollution is irreversible: the problem got "solved," but nobody truly understands it; knowledge didn't grow, only the answer count did. A theorem "solved" by a machine and understood by no one contributes zero to humanity's knowledge base — while doing very real damage to human curiosity.

The core accusation paraphrased in the AHM statement is a single sentence: publishing 700-plus files in one go isn't research at all; it's pure flexing of compute hegemony.

LeCun's camp: automated formal proof is mathematics' new era

On the other side, Turing Award winner Yann LeCun offered the exact opposite verdict: this, in his view, is substantive progress in automated formal proof — mathematics entering a new era. His logic is equally hard: a proof's ultimate value lies in correctness, and correctness is something machines can formally verify; "human readability" is just a byproduct of historical limitations, not the essence of mathematics. AI cracking 700-plus hard problems in one shot proves that the road of "machines understanding mathematical structure" is walkable — the rest is engineering. Demanding that machine proofs be readable by humans is like demanding cars look like carriages: refusing the new thing while pretending to welcome it.

Compress both positions into one sentence each: Tao defends the status of "humanly intelligible proof" as the carrier of knowledge; LeCun bets that "machine-verifiable results" will ultimately replace human proof. One side wants readability, the other wants problem-solving power. This isn't a flame war over who's right; it's a head-on collision of two philosophies of mathematics — and its echoes will reach the ears of everyone who writes code.

Here's a detail worth savoring that many will miss: Tao himself is an enthusiastic supporter of Lean, the mathematical formalization tool — years ago he led the effort to fully formalize the proof of a major conjecture in Lean. So his objection was never "machines shouldn't touch math," but "you can't skip the translation step." A formalized proof with a human-readable explanation is knowledge; 700 unexplained files dumped at once is noise. The real Tao-vs-LeCun disagreement isn't "machines or not" — it's whether machine output has to pass the "humans can understand it" gate before it counts. Remember that distinction; everything in my verdict hangs on it.

陶哲轩与杨立昆观点对立(示意图)

3. My verdict: they're not grading the same exam — but every vibe coder has to pick a side

The conclusion first: this debate will have no winner anytime soon, but both Tao's feared "pollution" and LeCun's promised "new era" will happen at the same time. A boycott statement won't bend the technology curve — OpenAI won't stop training models because mathematicians refuse to cooperate; if 8,000 problems can be brute-forced, so can 80,000. But the problem Tao points at is real and has no easy fix: once machines can mass-produce "correct but unreadable" results, the "understanding" link in humanity's knowledge chain gets bypassed. Mathematics is only the first battlefield; code, papers, design docs will all face this gate sooner or later.

And that is precisely the lived question of everyone in the vibe coding era — just under a different name: can humans still read, and dare to use, what AI generates?

Think about your own daily life. You let AI generate 3,000 lines of code overnight; all tests green, it runs. But do you dare merge it into the production branch? You let AI write a 40-page technical proposal — logically coherent, terminology on point — but do you dare sign it with a client? The reason you hesitate is the same reason mathematicians hesitate to use those 700-plus proofs: verifiable is not the same as intelligible; runnable is not the same as maintainable. A crack has opened between AI's "correctness" and human "comprehension," and it widens as models get stronger. Tests can tell you code "works," but only reading it tells you "why it works" — and the "why" is the entire basis on which you can reuse it, dare to change it, and debug it when things break.

One layer deeper: Tao's "academic pollution" has an exact twin in the engineering world called AI-accelerated tech debt. A machine "solved" a problem no human understood, so nobody can maintain it, nobody can improve it, nobody can debug it when it breaks — answers increased, knowledge didn't. Isn't that the exact smell everyone who's inherited AI-generated legacy code knows? What mathematicians smell now, and the suffocation you feel opening a 5,000-line AI-generated file, are the same odor. The difference is they wrote a joint statement about it, while you just cursed under your breath.

So my verdict, in three layers:

  1. Short term (1–2 years): the boycott is mostly posture; the technology won't stop. AHM's statement won't change OpenAI's roadmap, and individual mathematicians can hardly achieve a real "non-cooperation" — peer review, citations, collaboration networks are all entangled; a true cut would be academic suicide. But the posture itself has value: it forces the entire AI-for-Science community to answer the "readability" question head-on instead of pretending it doesn't exist. Sometimes asking the right question matters more than giving the right answer.
  2. Mid term: the real watershed is the "readability standard," not "problem-solving power." LeCun is right that formal verification will keep getting stronger; Tao is right that unreadable proofs can't enter humanity's body of knowledge. The future winners won't be the models that "solve fastest," but the ones that "can explain a proof like a human" — translating machine proofs into human-intelligible, teachable form will itself become a huge new lane. Same for vibe coding: code-generating AIs are everywhere; AIs that generate "code humans dare to merge" are the scarce ones. Next time you evaluate a coding agent, don't just look at its benchmark score — look at whether you'd put your name on the code it generates.
  3. Long term: mathematics may split in two. One kind is machine mathematics — formalized, verifiable, unreadable, running inside proof checkers; the other is human mathematics — intelligible, teachable, transmissible, running in papers and classrooms. The two connected by a "translation layer," the way compilers connect high-level languages to machine code today. That's not apocalypse; that's division of labor. What Scott Aaronson's "Mathocalypse" really prophesies is the death of the old division of labor, not the death of mathematics. When the old division dies, the "translators" in the new one will see their value skyrocket.

One last word for VibeFix readers. In this debate, both Tao and LeCun are right — but they're answering different exam papers. When it's your turn to hand in your paper, there's only one question on it: the next time AI hands you a "correct but unreadable" deliverable — code, a proposal, a proof, whatever — do you merge it straight in, or do you make it explain itself until you understand? Your answer decides whether you're a consumer of compute or an owner of knowledge. The mathematicians have already voted with a breakup; your vote is cast before every merge.

(Facts in this article are based on QbitAI's October 9, 2026 report and cross-reporting from CNA/CNYES and the DEV community; figures are as reported and marked "approximately"; quoted evaluations are paraphrases from the reporting, not verbatim originals.)

Views 0Comments 0

Comments (0)

ME
0/1000
Loading comments...

Sources

Browse projectsPublish your project

Related articles

Conceptual illustration of TianxiCode ranking first on the SWE-bench-Live leaderboard with a 71% solve rate
News
TianxiCode Tops SWE-bench-Live: Lenovo's Code Agent Hits 71% Solve Rate with DeepSeek-v4.1-Flash

Lenovo Tianxi AI's in-house code-agent framework TianxiCode, paired with DeepSeek-v4.1-Flash, topped the SWE-bench-Live Lite leaderboard at a 71% solve rate with official Verified certification. We unpack why this "real engineering" benchmark is harder, what "framework > model" really means, and three takeaways for vibe coding practitioners.

AI CodingProduct NewsIndustry Trends
DHH on stage at Rails World announcing 37signals has stopped writing code by hand
News
"We're Done Writing Code by Hand": Rails Creator DHH's Agent-Era Manifesto

Rails creator DHH announced at Rails World that 37signals is 'done writing code by hand.' This piece unpacks the November 2025 inflection point, the move to native apps and Rust, the counter-evidence from the same newsletter — and what vibe coders should actually take away.

AI CodingIndustry TrendsProduct News
Illustration of the Nemotron dual-gold recipe: SFT and RL checkpoints, 22000 curated programming problems, and the GenCorrect generate-evaluate-refine inference loop
News
NVIDIA Open-Sources Its 'Double Gold' Training Recipe: Nemotron Beats Top Human IOI Score, Full 22,000-Problem Dataset Released

NVIDIA's Nemotron systems hit gold level at IOI 2026 (535.4/600, above the top human score of 498.27 in an unofficial run) and IMO 2026 (30/42, graded by official IMO graders) — and the team open-sourced the full recipe: SFT/RL checkpoints, both training datasets, a new 200-problem olympiad benchmark, inference pipelines, and prompts. The lesson is co-design of model, data, and inference loop: GenCorrect's generate-evaluate-refine cycle carried a 291-point model past the 438.3 gold bar.

Open-source ProjectsModel UpdatesAI Coding