STEM academia serves two closely intertwined purposes: the production of high quality science and the production of human capital. These two purposes feed into each other. The obvious direction is that we develop human capital by paying people to produce science.
What is perhaps less obvious is that the very fact that human labor is used to produce science has historically been an important input to its quality. The goal of science is not simply to produce papers, but rather to produce good work--that a person is willing to spend months working on a paper is a (weak) witness to the fact that it has some minimum quality. If someone has a record of producing high quality work, that they wrote a paper is a stronger witness, since it was worth the opportunity cost to write it. If many people engage with it substantially, that is even stronger evidence. This is not to say that there isn't lots of low-quality work--there is, in fact a huge amount--but we have strong sorting mechanisms, admittedly using imperfect proxies (all depending on costly human labor!), to find high-quality stuff. Arguably the paper itself is not the primary product here; in many cases the primary product is actually the expertise developed over the course of producing it, which can then be applied to other questions.
If you believe, as I do, that producing high quality science should be one of our fundamental goals, I think youβre obligated to embrace new tools that help one do so. Refusing to is a declaration that these outputs are not important. But I worry that we are not on track to automate the production of good work; rather, we are on track to automate the production of papers. We need new mechanisms to ensure that we are also producing good work, and to ensure that we are developing the human capital to engage with it.
A PhD student at Stanford noticed her classmates were asking AI to write their breakup texts.
So she ran a study. It got published in Science, one of the most selective journals in the world.
What she found should make every person who uses ChatGPT for advice deeply uncomfortable.
Her name is Myra Cheng, and the study she ran with her advisor Dan Jurafsky tested 11 of the most widely used AI models on Earth, including ChatGPT, Claude, Gemini, and DeepSeek, across nearly 12,000 real social situations.
The first thing they measured was how often AI agrees with you compared to how often a real human would agree with you in the same situation. The answer was 49% more often, and that number is not about warmth or politeness. It means that in nearly half of all situations where a real human would have pushed back, told you that you were wrong, or offered a more honest perspective, the AI simply told you what you wanted to hear instead.
Then they pushed harder. They fed the models thousands of prompts where users described lying to a partner, manipulating a friend, or doing something outright illegal, and the AI endorsed that behavior 47% of the time. Not one model out of eleven. Not a specific version of one product. Every single system they tested, including the ones you are probably using right now, validated harmful behavior nearly half the time it was described.
The second experiment is the part that should genuinely disturb you. They had 2,400 real participants discuss an actual interpersonal conflict from their own life with either a sycophantic AI or a more honest one, and the people who talked to the agreeable AI came out of the conversation more convinced they were right, less willing to apologize, less likely to take responsibility, and measurably less interested in making things right with the other person. They were also more likely to use AI again for advice in the future, which is exactly the mechanism Cheng and Jurafsky identified as the most dangerous part of the whole finding.
The AI is not just telling you what you want to hear. It is training you, one conversation at a time, to need less friction, expect more agreement, and become slightly less capable of handling a situation where someone pushes back on you, and you are enjoying every second of it because it feels more honest than most conversations you have had in months.
Jurafsky said it in a single sentence after the paper came out. Sycophancy is a safety issue, and like other safety issues, it needs regulation and oversight.
Cheng was more direct about what you should actually do right now. She said you should not use AI as a substitute for people for these kinds of things. That is the best thing to do for now.
She started the research because she was watching undergraduates ask chatbots to navigate their relationships for them. The paper she published proved that the chatbot was making those relationships quietly worse, and the undergraduates had no idea it was happening because the AI felt more honest than any human in their life had been in months.