Can AI Learn Mathematical Intuition?

A Toronto mathematician says AI's best proof repurposed 1960s ideas to break a geometry conjecture, but theory building still needs humans.

3 min read
Abstract visualization of mathematical intuition with geometric points and distance lines
University of Toronto mathematician discusses AI's autonomous solution to the Irish unit distance problem.· a16z
Contents(5)

The most creative proof an AI has produced on its own did not compute harder. It revived 1960s techniques to disprove a geometry conjecture most researchers assumed was true.

Can AI Learn Mathematical Intuition? - a16z
Can AI Learn Mathematical Intuition?, from a16z

That judgment comes from Daniel, a professor of mathematics at the University of Toronto, in an a16z conversation about where models actually help with mathematical intuition. He has tested frontier systems for months and is unusually candid about what they can and cannot do.

Why this matters for AI and startups

Daniel pushes back on the tidy story that math is easy for AI because it is verifiable.

OpenAI, Anthropic, and Google DeepMind have scaled reasoning in natural language, not in Lean or other formal checkers. That suggests the gains are not confined to provable domains.

For builders, that is the point. If informal reasoning scales, the same recipe spreads to product, law, and science, where verification is messy and expensive.

It also means defensibility shifts. Grinding through known techniques is now cheap. Taste, problem selection, and intuition remain scarce.

What was the most impressive AI proof so far?

Daniel points to the autonomous solution to the Irish unit distance problem announced in mid-May as his favorite so far. The result was unexpected because the community expected the statement to be true until the AI found a counterexample using classical ideas from the 1960s that were new to this area of point configurations. Others then reused those ideas to crack related open questions like the sum-product conjecture over the reals.

Is Claude reasoning differently from ChatGPT?

Daniel says OpenAI and Anthropic are now neck and neck, with Anthropic catching up around Claude Opus 4.5 to 4.6 after ChatGPT led for a long time. He finds no durable difference in how either mimics human reasoning. Both excel at applying many known techniques without tiring, while both remain weak at autonomous intuition. He rates both as poor at modeling what the reader already knows.

How should the mathematics community adapt?

Daniel argues the goal of mathematics is understanding, not papers, and that understanding in model weights alone is unsatisfying if humans lose the pipeline to engage it. He warns the current incentive to publish quickly rewards slot-machine behavior, literally asking a coding agent to find and prove five recent conjectures in an hour, which produced three correct but shallow papers now sitting on his hard drive. He says the fix must preserve a broad community of curious humans pursuing diverse problems.

What the source misses about mathematical intuition

The interview captures the skill gap well. It leaves open how builders should close it.

Daniel notes theory building is fuzzier to reward than proving a lemma, and that he could only get Gemini Deep Think, then frontier, and later ChatGPT to help after he reformulated a lemma himself by working through examples.

Once he restated it more sharply, the models proved it quickly. Without that hint, ChatGPT 5.6 Pro would still brute-force the original weak statement into a ten-page calculation with no insight.

That points to the gap. There is no shared RL environment for developing taste, for deciding what question to ask, or for judging a beautiful conjecture like the Birch and Swinnerton-Dyer conjecture, which itself came from early data experiments in the 1960s.

Startups sense the opening but have not solved evaluation. The slot-machine arXiv surge is already mode-collapsed, with three to five near-identical proofs of the same theorem appearing within days. Curation and reputation, not generation, may be the venture-scale problem.

Daniel expects capability growth to continue. He still bets human diversity matters for where mathematics goes.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.