I find something strangely profound about how polar bears always have this look of boyish innocence to them no matter what the context is. Even covered in the blood of his prey, the bear looks like a silly little guy. Much to learn.
There’s a peculiar helplessness in watching the world race ahead in AI while being unable to participate at the frontier because of constraints completely outside one’s control.
I'm reflecting on how much research has changed since I've joined the PhD and wrote a short blog post about it (I joined in the tail end of the BERT era!). It seems pretty crazy how different processes are now, and I took the chance to do a retrospective before graduation:
https://t.co/OR7N8s4CgO
AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date.
However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001).
The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture.
With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104.
There is a lot of interesting commentary to be made:
1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date.
2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in!
3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig).
4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined.
Overall, an exciting development! Preprint is available here (https://t.co/YgiwgDF2qr) and will be on arxiv tonight; supporting code is here (https://t.co/KZhj15qDXC).
What if the model didn’t just use a computer, but actually was the computer?
Meta AI introduces "Neural Computer", a model where computation, memory, and I/O are all inside one learned system.
Their early prototype learns from screen recordings of terminals and desktops, and it can already imitate some basic computer behavior like rendering interfaces and responding to clicks or commands.
But it still breaks on slightly harder tasks like reliable reasoning, stable memory, and reusable skills.
I'm often asked how to land a research job at a frontier AI lab. It's hard, especially without a research background, but I like to point to @kellerjordan0 as an example showing it can be done.
Keller graduated from UCSD with no publication record and was working at an AI content moderation startup when he landed a cold call with @bneyshabur (who was at Google) and presented an idea to improve upon Behnam's recent paper. Behnam agreed to mentor him, which led to an ICLR paper.
Sadly there's less open research today, but improving upon a researcher's published work is a great way to demonstrate excellence to someone inside a lab and give them the conviction to advocate for an interview.
Later, Keller got on @OpenAI's radar thanks to the NanoGPT speed run he started. All his work was documented and it was easy to measure his success, so the case for hiring him was strong.
Keller is one example, but there's plenty of other success stories as well: 🧵
We are partnering with the Government of Tamil Nadu to build a Sovereign AI ecosystem in the State.
AI is becoming a core factor of productivity, shaping how effectively states educate their people, advance agriculture, improve healthcare, and deliver services to citizens at scale.
Tamil Nadu is the first state to invest in building a full-stack sovereign AI ecosystem, laying the foundations required to train, deploy, and embed intelligence across institutions and industries. Sarvam will build and train the models that power this ecosystem.
A Sovereign AI approach sets the right foundations from day one, enabling states and institutions to design, train, deploy, and govern AI systems that are safe, aligned, and rooted in their context.
We are excited to help bring this vision to life, and are grateful for the trust and support of the Government of Tamil Nadu. @trbrajaa@Guidance_TN
A new mind-blowing finding:
causal prediction alone is sufficient for strong visual learning, no need for any fancy reconstruction, masking, or contrastive loss
In this paper NEPA, instead of reconstructing pixels, they train a vision model to autoregressively predict the next patch embedding
and that alone yields strong visual understanding, matching or beating DINO and JEPA with a far simpler setup
now trending on alphaXiv 📈