Turns out I'm lousy at sticking to schedule. Dao AI 道愛 (an accessible book about programming and religion, AI language models and the trials, and tribulations, of using LLMs) is being delayed until it's ready - some time in the next month or so.
1st and 34th author of @GoogleDeepMind's paper [1] each got 1/4 Nobel Prize for protein structure prediction through Alphafold. Who invented that? (Disclaimer: a student from my lab co-founded DeepMind.)
The 2021 paper [1] failed to cite important prior work [2] by Baldi and Pollastri (2002): at a time when compute was roughly ten thousand times more expensive than in 2021, [2] introduced a pipeline very similar to the one of Alphafold 2, using multiple sequence alignment (MSA) to predict the secondary protein structure with the help of a position-specific scoring matrix (PSSM) or a profile matrix, going beyond even earlier work of 1988 [5][6][10]. The extra step (absent in Alphafold 2) was to predict the protein's topology, too. See also the follow-up work of 2012 [3].
[1] didn't cite @HochreiterSepp et al.'s first successful application [7] of deep learning to protein folding (2007, using LSTM instead of MSA to construct a profile).
[1] also failed to cite the essential prior work by Golkov et al (2016) [4][8], which had crucial aspects of AlphaFold: (1) identify homologous sequences in a database of proteins with known structure, (2) compute the co-evolution statistics using the homologous sequences, (3) train a graph NN to predict the protein contact map (that determines its 3D structure) directly from the co-evolution statistics, (4) demonstrate experimentally a significant boost in performance on the CASP dataset [4][9]. See the attached image!
Instead of the contact map, DeepMind (2021) predicted the distance map, and instead of graph CNNs, they used the quadratic Transformer published in 2017 (the unnormalized linear Transformer had existed since 1991 [11]). DeepMind also used more training data and much more compute for hyperparameter tuning etc.
Image credits: [4][8]
REFERENCES
[1] J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Zidek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli & D. Hassabis. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589, 2021.
[2] P. Baldi, G. Pollastri. A machine learning strategy for protein analysis. IEEE Intelligent Systems 17.2 (2002): 28-35.
[3] P. Di Lena, K. Nagata, and P. Baldi. Deep Architectures for Protein Contact Map Prediction. Bioinformatics, 28, 2449-2457, (2012).
[4] V. Golkov, M. J. Skwark, A. Golkov, A. Dosovitskiy, T. Brox, J. Meiler, D. Cremers (2016). Protein contact prediction from amino acid co-evolution using convolutional networks for graph-valued images. NeurIPS, Barcelona, 2016.
[5] N. Qian and T.J. Sejnowski (1988). Predicting the secondary structure of globular proteins using neural network models. J. Mol. Biol. 1988, 202, 865-884.
[6] H. Bohr, J. Bohr, S. Brunak, R.M.J. Cotterill, B. Lautrup, L. Norskov, O.H. Olsen, S.B. Petersen (1988). Protein secondary structure and homology by neural networks. The α-helices in rhodopsin. FEBS Lett. 1988, 241, 223-228.
[7] S. Hochreiter, M. Heusel, K. Obermayer. Fast model-based protein homology detection without alignment. Bioinformatics 23(14):1728-36, 2007. Successful application of deep learning to protein folding problems, through an LSTM that was orders of magnitude faster than competing methods.
[8] D. Cremers (July 2025). LinkedIn post on the Nobel Prize for AlphaFold.
[9] A Nobel Prize for Plagiarism. Technical Report IDSIA-24-24, 2024 (updated 2025) https://t.co/u9YxfBuqNf . Popular tweets on this:
https://t.co/heYSuPQDxp
https://t.co/QQU9FKpqAh
[10] The Nobel Committee for Chemistry (2024). Scientific Background to the Nobel Prize in Chemistry 2024.
[11] Annotated History of Modern AI and Deep Learning. Technical Report IDSIA-22-22, IDSIA, Switzerland, 2022 (updated 2025). Preprint https://t.co/YZrEphq1qx. This extends the 2015 award-winning deep learning survey in the journal "Neural Networks."
Universal Reasoning Model
Universal Transformers crush standard Transformers on reasoning tasks.
But why?
Prior work attributed the gains to elaborate architectural innovations like hierarchical designs and complex gating mechanisms.
But these researchers found a simpler explanation.
This new research demonstrates that the performance gains on ARC-AGI come primarily from two often-overlooked factors: recurrent inductive bias and strong nonlinearity.
Applying a single transformation repeatedly works far better than stacking distinct layers for reasoning tasks.
With only 4x parameters, a Universal Transformer achieves 40% pass@1 on ARC-AGI 1. Vanilla Transformers with 32x parameters score just 23.75%. Simply scaling depth or width in standard Transformers yields diminishing returns and can even degrade performance.
They introduce the Universal Reasoning Model (URM), which enhances this with two techniques. First, ConvSwiGLU adds a depthwise short convolution after the MLP expansion, injecting local token mixing into the nonlinear pathway. Second, Truncated Backpropagation Through Loops skips gradient computation for early recurrent iterations, stabilizing optimization.
Results: 53.8% pass@1 on ARC-AGI 1, up from 40% (TRM) and 34.4% (HRM). On ARC-AGI 2, URM reaches 16% pass@1, nearly tripling HRM and more than doubling TRM. Sudoku accuracy hits 77.6%.
Ablations:
- Removing short convolution drops pass@1 from 53.8% to 45.3%. Removing truncated backpropagation drops it to 40%.
- Replacing SwiGLU with simpler activations like ReLU tanks performance to 28.6%.
- Removing attention softmax entirely collapses accuracy to 2%.
The recurrent structure converts compute into effective depth. Standard Transformers spend FLOPs on redundant refinement in higher layers. Recurrent computation concentrates the same budget on iterative reasoning.
Complex reasoning benefits more from iterative computation than from scale. Small models with recurrent structure outperform large static models on tasks requiring multi-step abstraction.
You rarely solve hard problems in a flash of insight. It's more typically a slow, careful process of exploring a branching tree of possibilities. You must pause, backtrack, and weigh every alternative.
You can't fully do this in your head, because your working memory is too limited. Writing is the external medium that affords the time and precision necessary.
Serious thinking must be done in writing. And that's why you can't outsource your writing, because then you're outsourcing your thinking.
Interesting research from Google.
Research has shown that neural networks don't just memorize facts. They build internal maps of how those facts relate to each other.
The view of how transformers store knowledge is associative: co-occurring entities get stored in a weight matrix, like a lookup table. The embeddings themselves are arbitrary.
But this view can't explain something these researchers found.
This new research demonstrates that transformers learn implicit multi-hop reasoning when graph edges are stored in weights, even on adversarially-designed tasks where associative memory should fail. On path-star graphs with 50,000 nodes and 10-hop paths, models achieve 100% accuracy on unseen paths.
This geometric view of memory challenges foundational assumptions in knowledge acquisition, capacity, editing, and unlearning. If models encode global relationships implicitly, it could enable combinational creativity but also impose limits on precise knowledge manipulation.
Paper: https://t.co/Rk68BdRRcG
Learn to build effective AI agents in our academy: https://t.co/zQXQt0Pem8
A beer-based vaccine has been developed in the United States — virologist Chris Buck brewed beer capable of protecting against diseases
He managed to integrate virus-like particles into yeast, helping the body fight infections.
He and his family tasted the “foamy shot,” and the results were striking: after drinking five liters of the beer, their bodies really did produce antibodies that could help fight illness.
Now he’s considering how to make beer-vaccines for other diseases as well.
I remember seeing a post about how normalized AI in advertisements making the company look cheap will mean real art will start becoming the standard for brands to seem more luxurious and it's really happening
The single most revolutionary idea in my economic education came from reading Harvard Professor Michael Porter and realizing that pure free-market competition is for absolute chumps.
If you are a business, free market competition is your sworn enemy. It is your imperative to subvert it at every turn.
Unconstrained competition is a corrosive acid, that left unchecked, relentlessly erodes profit.
Therefore, the duty of any business is to do at least one of five things:
- Become the dominant seller to its buyers
- Become the dominant buyer from its sellers
- Create a product with few, if any, substitutes
- Construct, or have constructed for it, persistent barriers to new entrants
- Minimize competition through market structure or tacit coordination
The patent system, which we treat as entirely natural, is just the third point: a state-granted exclusion right that suppresses the possibility of substitutes.
If a business has none of these advantages, excess returns do not persist.
They are competed away until they converge to the absolute floor of sustained profitability, which is the risk-adjusted return on passive investment.
Capitalists often sell it as if the businessman, and chief among them, the billionaire is the natural ally of the consumer. They are not.
It is their job to evade constraint, and it is our job to constrain them.
not getting into a philosophical debate, but this book really changed how I see the topic and made me feel more humble. human intelligence is impressive, but calling it ‘general’ isn’t very objective. my cat would disagree.
to me human intelligence is better seen as socially driven cognitive adaptations, and there’s a huge WORLD of intelligence we still don’t understand, and are nowhere near recreating with current AI
Six authors are suing Anthropic, Google, OpenAI, Meta, xAI and Perplexity for what they call “a straightforward and deliberate act of theft that constitutes copyright infringement.”
The lawsuits over AI training show no signs of abating.
https://t.co/SWZNEH2Ocf
A massive discovery in neuroscience: fMRI signals don't always match true neural activity. In ~40% of cases, signals increased where activity actually decreased. This challenges the core assumptions of tens of thousands of studies.
#Neuroscience#fMRI#BrainResearch #CognitiveScience #MedicalResearch https://t.co/jhZ9Rluq5X via @neurosciencenew
People crying that this hinders artificial intelligence (AI) development don't understand that this creates an evolutionary pressure for AI to become truly intelligent, to not depend on tons of high quality work.
Only lazy & shady tech companies sees this as an obstacle.
🚨 The UK government just published a breakdown of the responses to its consultation on AI & copyright:
- 95% of respondents want AI companies to pay for their training data (made up of 88% saying strengthen copyright law, & 7% saying leave it as is)
- Only 3.5% want a new copyright exception for AI training (3% want an exception + opt-out for rights holders, 0.5% an exception with no opt-out)
These results are absolutely overwhelming. The government should rule out a new copyright exception immediately.
So, I grew up in total poverty in a farming town in southern Texas. By a lucky chance I was able to attend university. The first week on campus I set foot in the student library, and what followed was one of the most transformative experiences of my life. (thread)
Back in 2019, ARC 1 had one goal: to focus the attention of AI researchers towards the biggest bottleneck on the way to generality, the ability to adapt to novelty on the fly, which was entirely missing from the legacy deep learning paradigm.
Six years later, the field has responded. With test-time adaptation, we finally have reasoning models capable of genuine fluid intelligence.
While ARC 1 is now saturating, SotA models are not yet human-level on an efficiency basis. Meanwhile ARC 2 remains largely unsaturated, showing these models are still operating far below the upper bound of human-level fluid intelligence. We're still only at a fraction of what a human mind is capable of in a single sitting with no external tooling (a level which is itself significantly above a full score on ARC 2), so there's more work to be done.
And as we get closer to AGI, the challenge goes beyond fluid intelligence. The new bottlenecks are exploration, goal-setting, and interactive planning. We are releasing ARC 3 in Q1 2026 to target exactly this. It's time to trigger a new class of breakthroughs.
To perfectly understand a phenomenon is to perfectly compress it, to have a model of it that cannot be made any simpler.
If a DL model requires millions parameters to model something that can be described by a differential equation of three terms, it has not really understood it, it has merely cached the data.
This place is toxic.
For the last seven years I warned you that LLMs and similar approaches would not lead us to AGI. Almost nobody is willing to acknowledge that, even though so many of you gave me endless grief about it at the time.
I also warned you -– first –- that Sam Altman could not be trusted, that OpenAI would lose its dominance, that GPT-5 would not be all AGI, and that LLMs lacked world models. That hallucinations would not go away. That out of distribution generalization was THE key issue.
And that the economics of LLMs didn't make sense. And that the LLM companies would start seeking bailouts.
The receipts are all here if you care. As for me, I have had it. If you want to hear other prescient warnings in advance, subscribe to my newsletter Marcus on AI.
Or you can stay here and be lied to; the choice is yours.
Looking forward to maximal AI research space exploration in the future. Just a shame that it's taken a giant speculative market bubble, the pending collapse of the US economy and destruction of Gen Z's jobs and social contract to get there.
We just don't learn do we...
@nirsd@DJFreshUK 100%, and honestly it would be a refreshing change. Everyone, and their pet dog, in the AI space is all over the current paradigm. I think it would be beneficial for more diverse research to be explored.
Fascinating how all the propaganda accounts turned out to be visibly based out of US adversary countries -- you'd think their intelligence services would be better at maintaining US proxies. I guess they realized they didn't need to care about hiding
The reason new technology often gets released as a toy first is that it has to go from "doesn't work at all" to "works perfectly", and the first step along the way is "works a tiny bit, not enough to be practically useful, but hey it's new and kinda cool"
Always important to see new tech as a point along this trajectory, not as the end result