It’s hard to overstate how much AI has changed my research life. I don’t know a single PhD student who doesn’t use it. I probably don’t know a single professor who doesn’t use it, either. Friends who swore these models were useless for math six months ago now all have pro subscriptions, sometimes to multiple models, and are happily paying hundreds of dollars per month for access.
It’s undeniable that these systems are better at proving and disproving theorems than most of us. Maybe this stops as you approach the outskirts of the most abstract subfields of math, but for most things that a PhD student like me is encountering, AI is simply miles ahead of you. One can quibble about whether this ability is due to more “true insight” or simply more memory, but the effect is the same.
My sense is that most researchers are uneasy about this new AI era. Despite being able to prove more things than ever, faster than ever, there is a sense that something is being lost.
It’s tempting to chalk this concern up to the fact that the job of a researcher is changing, and changing quite substantially. But this isn’t quite right. We’ve seen radical changes previously, and from what I can tell, those changes were unaccompanied by any anxiety. Before the computer, a large part of a mathematician’s job was doing calculations. After the computer, there wasn’t much complaining about the fact that logarithm tables now didn’t need to be computed by hand. It was clear that removing the need to spend time on rote arithmetic freed up a mathematician to spend more time on the kind of research that mattered.
So too with statistical software, which enabled analyses to be run much faster, and with search engines that allowed us to quickly find new papers and preprints instead of waiting for journal proceedings. These had a huge effect on math and statistics—and science more generally—but no one felt the job was being threatened.
Can we say the same about AI? Maybe. We might hope that AI can automate the mechanical parts of proving theorems, leaving more time for theorizing, developing intuition, and asking interesting questions.
Unfortunately, I think this is wrong. You can’t decouple proving theorems from the rest. Math isn’t cleanly split into the “proving theorems” part and the “creative vision” part. Proving theorems is the laborious part of doing math, but it’s also necessary for building the knowledge required to ask fruitful questions.
In the old days (i.e., five months ago), you might have spent a summer working on several problems you thought were interesting and important. Now, once the problem is well-formulated, you might get an answer from an LLM in an afternoon. Your job is to understand the answer and present it in a digestible way to the community.
This can feel rather deflating. Like checking a student’s work, correcting it on the margins, and maybe simplifying it a bit. But ultimately, it doesn’t feel like yours. And it isn’t yours! Beyond being subjectively disorienting, it also poses a problem for progress in math. The process of spending your summer playing around with ideas, pursuing the wrong direction, deeply understanding the problem, was how you learned what was important, what question to ask next, and why the community had converged on certain techniques.
In other words, proof in math is not an afterthought. It’s part of the process by which you decide whether what you’re working on is interesting, develop intuitions, and work out whether your initial question was worth answering in the first place. Being handed the answer short circuits this process.
As usual, Terence Tao got here first and said it better than I could. He writes:
Working out whether a question is actually worth highlighting is a lengthy, deliberate, and subjective process, often informed by historical experience on what good mathematics was generated (or not generated) while working on earlier problems of this type. In particular, being aware of the “difficulty landscape” in a field - what questions are very easy to answer with known methods, which ones can be solved but only with some effort, and which ones are impossible - is of crucial importance in making such determinations.
He goes on:
In short, the indiscriminate use of powerful solution-extraction tools can achieve the immediate short-term goal of solving problems at hand, but at the cost of sustaining the ecosystem for the next wave of progress, or in understanding the progress already obtained.
So we’re proving more things than ever, but there’s an open question as to whether we’re understanding more things than ever. The understanding used to come from locking yourself in a room for several months and banging your head against a wall, which we no longer do.
With AI, every researcher is instantly promoted to principal investigator. The research ladder used to be something like this: as a PhD student you spent most of your time writing your own papers. As you progressed through academia, you spent more and more time advising students on the questions that you thought were promising. You spent less time in the messy details, and more time coordinating efforts at a higher level and carrying out a research vision. Concretely, this meant you had a lab with several students (more or less depending on the field), and you were advising the students on what to study.
Importantly, these senior researchers had built up research taste by spending many years as students and early-career researchers in turn. They had spent a long time banging their heads against the wall. Access to AI is like letting everyone have access to tens, hundreds, or thousands of technically brilliant PhD students who can work for them 24/7 on whatever problems they choose, without necessarily having built up the research taste about what problems are worth solving.
To be clear, this isn’t a worry about gatekeeping, or of trying to maintain established academic hierarchies. But it is a worry about the quality of questions that are being asked and answered. There are an infinite number of possible questions. Many of them are boring, unenlightening, already known, trivial, or too hard. Knowing what question is a good one requires knowledge, and this knowledge comes about by interacting with the messy details of a field. In the case of math, this means writing and struggling with proofs.
And importantly, this knowledge used to be idiosyncratic. Different researchers had different implicit (and explicit) knowledge from working on a particular set of questions with a particular set of techniques, which meant they would explore different parts of the question space than their neighbors did.
If you’re not building up this idiosyncratic knowledge then you become fungible. This seems to be a common feeling among PhD students these days. If anyone else can ask the same question to an LLM, and get as good of an answer as the one I would have given, what am I really adding? Is it important that I’m doing research in addition to my office mate, or should we just buy him two subscriptions to ChatGPT Pro and let him cook?
I realize I sound quite pessimistic, which is unusual for me. I should say that, ultimately, I think this is a solvable problem. And I even suspect that the solution comes from using these tools more and better, not from using them less, or putting any restrictions on who can use them. But I do think that the way many of us are using them now is prioritizing short-term progress over long-term understanding.




With the capabilities of AI, theoretical research should be called “interesting” iff someone outside the department finds it interesting. My understanding of the history of mathematics is that this was once the case. Recently mathematics has become self-referential, with the defense that “discoveries in pure math have applications down the line”. This was true, but seems increasingly irrelevant if AI can solve any well-defined problem. I’m more optimistic that the AI capabilities can reset the focus on what is “interesting” (e.g. Tsimerman working on AI safety).
There is a different question here on what to do with education: there is agreement for undergrads and below, but where does a PhD student land? I think there should be a place for a PhD student who has 20 lean-certified theorems that are used in academia and industry, but does not understand the technical details of any of them.
The question is if math is a science or an art. The mathematical artists will need to up their game to make results more presentable and appealing (we can think of these researchers as the "Apple" researchers who don't do breakthrough first but do it the best). But we should not eschew the scientist that wants to enact progress using the field instead of contributing to it.
Ben, nice post!
Two comments: ages ago I incidentally wrote an undergrad piece on a related problem (it's not very good and I'm not going to link it here. I even think it's in Portuguese, in any case): math-assisted proofs and what that meant to the epistemology of mathematics. I was thinking of mechanically generated proofs, of course. If we define a proof as a sequence where everything is either the result of applying some derivation rule to previous steps or an axiom, there is no real difference between brute-forced proofs and 'elegant' ones. I looked up a bit into the mathematical community's answer to the 4-color theorem, the first major problem solved by computers. I think it points to that notion of proof being insufficient and mathematical research being a social process where fuzzier notions like 'insight' and 'explainability' play a larger role than that admitted by the formal paradigm. In a sense, this is different from LLM-generated proofs, that resemble the human-insight procedure more closely. They are not simply brute-forcing combinatorics. But this difference may be one of degree: Tao's comment highlights some of this and Buckmaster's comments on how he was working with the LLM-generated arguments point to a slight weaker notion of proof. Of solving statements without gaining the same degree of understanding, whatever it may be.
Second one: there's a Ted Chiang short story about a world in which humans are overcome by AIs in science and human science becomes just trying to inspect the AI-generated corpus. Almost like archaeologists to discover things about the universe. In a way, digesting and interpreting proofs, and trying to understand them, and presenting them to the community, is eerily similar to Chiang's story. Pardon any misquote, I read it a decade ago and am citing from memory, but it's food for thought in any case.
But I'm really interested in your positive propositions to 'solve the problem'. Writing to encourage you to write them down at some point.