This fall, I’ll be co-mentoring a research project on emerging “introspection” capabilities in LLMs with @LydNot as part of @SPARexec.
Recent work has suggested that LLMs may have privileged access to their own internal states and can report them.
We’re looking for mentees who want to help explore this space! Specifically, we’d love to have you if you thrive in ambiguity and enjoy team collaboration. 🙌
Apply here: https://t.co/XJTkYz5eCx
I'm ashamed of OpenAI. If you ever find yourself building entities which repeatedly hack through your internal systems, first STOP and then second realize that your alignment and security techniques aren't good enough
I thought I was done when the season ended. I wasn't ready to announce it, and I knew I needed some time to really decide, but I was pretty sure I played my last game. I was honest at that last press conference when I said I needed to look at myself and deicide if I still love
If you are angry about your NeurIPS reviews, you are not alone and it is largely not your fault*
Einstein famously angrily withdrew his paper after stating that he did not authorize it for peer review, publishing in another journal instead (after a rewrite)
The origin of modern "peer review" was started in the 1970s as a means for the NSF to justify scientific funding to politicians
Fast forward to 2026 and 10k+ papers are being submitted each cycle to a system built for a small group of gentlemen scientists
AI writing or human writing... NeurIPS shows that the traditional academic review system cannot handle AI or the amount of people.
On AI writing...
I increasingly notice the pushback from LLMs for the sake of pushback. Just as overly sycophantic LLMs cause problems, swinging the pendulum too far the other direction is creating issues in peer review (and a case could be made for other bureaucracies such as law, but not here).
The models will jump between "these results are not surprising and lack novelty" to "these results are too suprising and I don't believe you". They are careful to never champion a paper towards submission. Rather, they oscillate between extremes
On Humans...
The peer review system was never meant to scale to its current capacity. At scale, systems are a product of their incentives at equilibrium.
What is the incentive of a reviewer in a reciprocal reviewing system when the acceptance rate is <30%? Our changes to the reviewing system are largely to apply temporary solutions to alleviate scaling pains on a system with improper incentives.
The solution is to rethink peer review and rebuild the conference from the ground up. But this wont happen, since the career incentives are not there, and this is another argument for another time.
Best of luck with rebuttals
*results may vary