Last time I confessed some worries I had about how AI is affecting research in math and statistics. I ended by saying that I was ultimately optimistic, but that it would take developing new norms and practices about how to work with these tools effectively.
Luckily, I recently got a glimpse of what these norms and practices might look like. They came from a lunch I had with Nikita Zhivotovskiy, a faculty member at Berkeley.
My sense is that most people are using AI to do the same things they would have done otherwise, but faster. They have a question they’d like to answer. They ask AI. AI answers it. They read the answer, go through it to make sure it’s correct (hopefully), write a paper about it, and then repeat.
This process is helpful in the short term—we’re answering our questions after all!—but my worry is that it’s harmful in the long term, both individually and collectively. As anyone who has ever learned anything ever can attest, checking if you agree with someone else’s answer to a question is different from deeply understanding the question yourself.
Before LLMs, deep understanding was a byproduct of doing research and writing good papers (emphasis on the good), since proposing a coherent and novel solution to a problem required real insight. But now we can write papers with a surface-level understanding of both the problem and its solution. This hurts our ability to develop intuition and “research taste,” that ever-elusive sense of what is worth focusing on in an infinite sea of possible questions.
Talking with Nikita made me realize that one need not use AI this way. And the alternative is not to forgo using AI. It’s to use AI more and to use AI better. It’s to use AI to obsessively help forge your own understanding.
This sounds mundane and obvious. “Don’t use AI poorly. Use it well. Use AI to help you understand things better and develop intuition.” You don’t say! But I’m convinced that many people are not actually doing this. Nikita was doing it.
What does this look like concretely? Here are some Nikita-inspired ideas.
In math and stats specifically, you can now iterate on proofs tens to hundreds of times before you’re satisfied with them. If AI generates a proof of something, first have 5-10 independent agents try to find flaws in it. Once there’s agreement about correctness, you can boil it down to its essence. What are the key steps in the proof? What is the path through idea space? What objects are being used, and how?
Write down those steps—the intuition, not the details—for yourself, then have independent agents try to recreate the proof from only that outline. Can they do it? Have you missed a key step? Do any of the agents take another path and omit some of the steps? Do any of them generate a significantly shorter proof? If you don’t understand why the proof relies on doing things a certain way, try to have an agent generate a proof using a different method that you suspect might work. Why does it work or not work?
You now have the opportunity to not just have AI answer questions, but to better understand the landscape of possible answers. You do this by continuously probing the proof, understanding what steps are crucial and why.

Once you start thinking in terms of a proof landscape, it’s easy to think of more techniques. For instance, since simulations are now cheap to set up and run, you can simulate inequalities to see if there’s any discrepancy between the analytical bound and what you’re losing experimentally. You can sweep the parameter space to figure out precisely for which parameters the theorem breaks. The right proof should shed light on exactly why it fails at those values.
And others: For each assumption in the theorem, you can have agents try to prove the result without that assumption. Conversely, given the steps of a proof, ask them to reconstruct the statement that’s being proved. Is it different from the original statement?
Notice two things about all of these techniques: How central the human still is in the research loop, and also how much AI is being used. Naive strategies try to replace human labour with AI labour, and there’s a zero-sum competition between who is doing what. But if you’re using these tools thoughtfully, human labour and AI labour are not in tension.
Also note that none of this depends on these tools plateauing at their current capabilities. Regardless of how good they are, insofar as the ultimate goal of research is human understanding, it will always take lots of iteration and prompting for the human to fully understand the problem and its solution.
Of course, all of this assumes that the researcher is trying to write better papers and increase their own understanding. Whether researchers want to do this, or will be incentivized to do this, are different questions. Maybe we get sucked into a vortex of writing thousands of technically-correct-totally-uninteresting papers so that our h-index increases. (Judging by the quality of recent submissions to conferences, we seem to be veering in this direction.)
I think we’ll go through a period of turbulence as low-hanging fruit gets quickly picked by humans pointing AI at open problems, peer review buckles ever more, and departments sort out what to do with increasingly desperate and despondent grad students. But I’m confident that there is a world in which these tools help us do even deeper research than was previously possible, even if not everyone opts into it.


