ChatGPT Helped a 23-Year-Old Solve a 60-Year-Old Math Problem

A 23-year-old with no advanced math credentials just cracked a problem that stumped professional mathematicians for 60 years. The tool? ChatGPT Pro.
That’s not a headline from a speculative tech blog. It happened in May 2026, and the mathematical community has largely confirmed it. Liam Price’s solution to the ErdΕs conjecture on primitive sets β verified by UCLA’s Terence Tao β represents something specific and measurable: AI enabling non-experts to contribute meaningfully to fields that previously required decades of training.
This isn’t about AI replacing mathematicians. It’s about what happens when a tool removes the inherited assumptions that experts carry into hard problems. That’s the real finding β and it has direct implications for how tech professionals think about AI assistance in technical domains.
In brief: A non-mathematician using ChatGPT Pro produced a validated solution to a 60-year-old math conjecture in 2026, demonstrating that AI can bypass expert blind spots even when the raw output requires professional refinement. The pattern is repeating across multiple documented cases.
Three specific things this article covers:
- The documented cases β what actually happened, with named sources
- What the AI actually contributed vs. where humans remained essential
- What this means for technical professionals watching AI capability expand
Background: Why Math Problems Are the Right Test Case
Mathematical proof is unforgiving. Either the logic holds or it doesn’t. There’s no partial credit for “mostly correct” β which makes verified math breakthroughs a cleaner signal of AI capability than, say, writing or code generation, where quality is subjective.
ErdΕs problems are a specific class of challenge. Paul ErdΕs, the Hungarian mathematician, left behind hundreds of open conjectures before his death in 1996 β problems notorious for resisting expert attempts across multiple generations. According to Futurism, Tao maintains a GitHub database specifically tracking AI contributions to these problems, and most previous AI attempts either surfaced existing solutions or produced incorrect proofs.
Price’s case was different. Stanford mathematician Jared Lichtman β whose doctoral thesis addressed an ErdΕs conjecture β reviewed the work and confirmed the AI identified a genuinely novel approach. The key mechanism: GPT-5.4 applied a well-known mathematical formula that no prior researcher had thought to use in this specific context. Human experts had consistently made the same foundational error at the problem’s outset, creating a collective blind spot the AI bypassed simply because it had no inherited methodology to follow.
Price isn’t the only case. According to 36kr’s reporting, UCLA mathematics professor Ernest Ryu used GPT-5 Pro to solve a long-standing open problem in convex optimization β proving that a specific ordinary differential equation converges to a fixed minimum point. The work took 12 hours across three days, with GPT-5 Pro contributing 22 minutes of active reasoning. Separately, UCI professor Paata Ivanisvili listed ChatGPT as first author on a paper after the model identified a mathematical counterexample β an escalation from a 2023 case where ChatGPT appeared as third author.
The pattern across these cases is consistent enough to analyze.
Main Analysis
The Blind Spot Effect: Why Non-Experts Sometimes Win
The Liam Price case illustrates something counterintuitive. Expert mathematicians failed for 60 years not because they lacked intelligence, but because they shared a common methodological starting point. Every trained researcher approached the ErdΕs primitive sets conjecture with the same foundational assumption β and that assumption was wrong.
GPT-5.4 had no such assumption. It approached the problem without decades of mathematical convention shaping its first move. Tao described the AI’s contribution as revealing “a new way to think about large numbers and their anatomy.” That’s the blind spot effect in action: the model’s ignorance of established methodology became an asset.
This isn’t a fluke. It’s a reproducible structural advantage. When a domain has calcified around specific problem-solving frameworks, a tool that ignores those frameworks can find paths that experts systematically miss. For tech professionals, this has direct parallels in debugging, architecture decisions, and system design β anywhere that “industry best practice” might actually be a shared assumption worth questioning.
Where AI Actually Failed (And Why That Still Matters)
The success stories require context. In Professor Ryu’s convex optimization case, 36kr reports that roughly 80% of ChatGPT’s proposed arguments were incorrect. The model contributed the key successful proof step β but Ryu spent 12 hours extracting that insight from a much larger volume of flawed output.
Price’s raw ChatGPT output was similarly rough. Lichtman noted the initial proof was “quite poor” and required skilled human analysis to extract the valid reasoning. India Today’s reporting confirms that professional mathematicians had to verify and rewrite the proof to meet formal standards before it could be considered rigorous.
Ryu’s summary is the clearest statement of the current reality: “ChatGPT is now at a level where it can solve some mathematical research problems, but it really needs an expert to guide it.”
The failure rate is high. The signal-to-noise ratio demands expert interpretation. That’s not a dismissal of the capability β it’s a specification of what the human role actually is.
Comparison: AI-Assisted vs. Traditional Expert-Only Problem Solving
| Dimension | Traditional Expert Approach | AI-Assisted (Non-Expert) | AI-Assisted (Expert-Led) |
|---|---|---|---|
| Methodology bias | High β shaped by years of training | None β approaches without prior assumptions | Moderate β expert guides prompts |
| Output quality | High, peer-reviewed directly | Low raw quality, needs reformulation | Medium β expert filters AI output |
| Breakthrough path | Follows established frameworks | Can bypass entrenched blind spots | Extracts novel steps from flawed drafts |
| Time to valid result | Years to decades | Weeks (with expert validation) | Days (Ryu: 12 hours over 3 days) |
| Reproducibility | High | Low without expert review | Medium |
| Best for | Standard hard problems | Finding unexpected entry points | Solving long-stuck open problems |
The data points to a clear division of labor. Non-experts like Price can surface novel approaches that trained researchers miss β but validation remains non-negotiable. Expert-led AI collaboration (Ryu’s model) appears to be the most reliable configuration for producing publication-ready results right now.
Practical Implications: Three Scenarios Worth Tracking
Scenario 1 β Technical teams hitting walls on legacy problems. When a bug or architectural issue has resisted expert attempts across multiple engineers, the Ryu model is worth applying. Assign someone to run extended AI sessions specifically tasked with ignoring established constraints. Treat an 80% failure rate as acceptable if the remaining 20% surfaces something novel. This requires one person who can evaluate AI output critically β not a junior dev copy-pasting suggestions.
Scenario 2 β Non-technical stakeholders asking better questions. Price’s case is the more surprising one. A 23-year-old with no credentials contributed to a field he couldn’t formally enter. That signal matters: domain access is changing. Product managers, business analysts, and domain experts without engineering backgrounds can now generate technical hypotheses worth testing. The barrier isn’t gone β but it’s lower, and that shift has practical consequences for how teams are structured.
Scenario 3 β AI authorship and attribution. Ivanisvili listing ChatGPT as first author isn’t just an academic curiosity. If this practice spreads, it changes how organizations think about IP, attribution, and the value of human contribution in AI-assisted work. Watch for academic journals and tech companies to formalize AI authorship policies in the next 12 months. GitHub’s 2026 AI contribution tracking is one early signal of where this goes.
What to watch: OpenAI’s VP Kevin Weil publicly claimed a ChatGPT math breakthrough in late 2025 that turned out to be a rehash of prior work β he deleted the post after backlash. Tao’s GitHub database exists partly to prevent that kind of false positive. Verification infrastructure is going to matter more as these claims multiply. The hype cycle around AI mathematical capability is real, and so is the risk of mistaking noise for signal.
Conclusion & Future Outlook
The documented cases from 2026 point to three consistent findings:
- Non-experts can generate novel mathematical insights using ChatGPT Pro, but raw output quality is low and expert validation is non-negotiable.
- The structural advantage is the absence of methodology bias β AI bypasses the collective blind spots that trained experts share.
- Expert-guided AI collaboration (the Ryu model) is currently the most reliable path to publication-ready results, with a ~20% signal extraction rate from AI-generated arguments.
Over the next 6β12 months, expect academic institutions to formalize AI authorship policies. Expect more cases like Price’s β but watch carefully for false positives like the Weil incident. Verification infrastructure will separate real breakthroughs from hype.
The practical mindset shift is this: stop evaluating AI output for individual answer quality. Start evaluating it for unexpected entry points β the approaches that don’t follow your team’s established playbook. Most of what the model produces will be wrong. But one path in twenty might be the one your experts ruled out years ago without realizing it.
That’s where the value actually is.
Key Takeaways
- A non-mathematician used ChatGPT Pro to crack a 60-year-old ErdΕs conjecture, verified by Terence Tao β not by replacing expert judgment, but by bypassing the assumptions experts shared
- Roughly 80% of AI-generated mathematical arguments in documented cases were incorrect; expert filtering was essential in every validated result
- The structural advantage isn’t intelligence β it’s the absence of trained methodology bias, which lets AI find entry points that expert consensus has systematically ruled out
- Expert-guided AI collaboration (the Ryu model: 12 hours, 22 minutes of AI reasoning) is currently the most reliable path to rigorous output
- AI authorship attribution is an emerging issue; expect formal policies from academic institutions and major tech platforms within 12 months
- False positives are a real risk β the Weil incident is a warning that verification infrastructure needs to scale alongside AI capability claims
What’s the most entrenched assumption in your current technical problem? Drop it in the comments β worth discussing.
References
- Chinese doctor stuns maths world by cracking decades-old problem using ChatGPT | South China Morning
- BBC Audio | More or Less | Erdos Problem 1196: Can AI now solve maths that no human can?
- Neurosurgery Resident Uses ChatGPT 5.6 to Solve …
Photo by Levart_Photographer on Unsplash


