For Joel David Hamkins, the problem is not only errors, but the machine’s insistence on appearing correct even when it is wrong.
Artificial intelligence has been presented as a powerful ally of science. There are reports of systems that have allegedly helped solve complex mathematical problems, including challenges related to the work of Paul Erdős. However, when these tools are put to the test by practicing mathematicians, the enthusiasm does not always hold up.
It was precisely this gap that led Joel David Hamkins, a mathematician and professor of logic at the University of Notre Dame, in the United States, to share a direct assessment of the current use of AI in mathematics. In an interview with the Lex Fridman Podcast on December 31, Hamkins described his hands-on experience with large language models — and the results fell far short of expectations.

A mathematician testing AI in practice
Hamkins does not speak from future projections or abstract tests. He reports having experimented with different AI systems, including paid models, applying them to real mathematical questions. His conclusion was straightforward:
“I’ve tried it and tested it, but I didn’t find anything useful. Basically, zero. It doesn’t help me at all.”
Even so, he makes an important distinction right at the outset:
“I think I would make a distinction between what we have currently and what may emerge in the coming years.”
In other words, his criticism refers to the current state of the technology, not a definitive rejection of AI’s potential.
Where AI fails, according to Hamkins
The central problem, for Hamkins, is not simply that AI makes mistakes — something expected of any tool — but the kind of mistakes it makes. According to him, when dealing with mathematical questions:
“My typical experience is that the AI gives answers that are nonsense, that are not mathematically correct.”
In advanced mathematics, plausible but incorrect answers are especially problematic. Every argument must be logically consistent and verifiable, something current models still cannot reliably guarantee.
The most delicate point: resistance to correction
Beyond the errors themselves, Hamkins draws attention to how models handle criticism. He describes interactions in which he points out specific flaws in the reasoning presented, yet does not see consistent correction from the system.
“What’s frustrating is when you have to argue about whether the argument is correct or not, you point out exactly the error, and the response is: ‘Oh, it’s fine.’”
According to Hamkins, this resistance to correction makes it difficult to build a productive dialogue, something essential in collaborative mathematical work. He notes that, in a human context, this type of interaction would tend to be avoided:
“If I were having an experience like that with someone, I would simply refuse to talk to that person again.”
A criticism shared by other mathematicians
Hamkins’ assessment is not isolated. Mathematicians such as Terrance Tao have also pointed out that AI systems can produce proofs that appear well structured but contain subtle errors, difficult to detect without careful analysis.
The problem, therefore, is not just making mistakes, but making them with excessive confidence and resisting feedback — something that undermines the trust required for collaborative mathematical work.
Between impressive tests and real-world usefulness
Hamkins also highlights the difference between strong performance on benchmarks and practical usefulness in scientific research. High scores on standardized tests do not necessarily mean that AI is ready to act as a reliable partner in high-rigor mathematics.
Between advances and limits of AI in mathematics
Joel David Hamkins’ experience suggests that, despite recent advances, current artificial intelligence models still face significant limitations in advanced mathematics. Frequent errors, difficulty correcting them, and overly confident answers reduce their usefulness in formal research contexts.
His position is not one of definitive rejection, but of skepticism grounded in practice. For Hamkins, the future of AI in mathematics remains open — but the present, at least for now, still calls for caution.