Extraordinary claims require extraordinary evidence so check out our release blog for more technical info: https://t.co/wR5iMBbKeq
Join our waitlist for early access: https://t.co/DXXqUCF0nz
Have technical chats and meme with us on Discord (rumors are good memers skip the line): https://t.co/cAIll7nGIe
We are so pumped this is in the hands of developers now, and we are just getting started!!!
DEEPSEEK ACABA DE MATAR TODA LA INDUSTRIA DEL CODING AGENT
Se llama deepseek-harness
Es el framework mas completo para crear agentes de codigo
Open source. El plan mas completo de Claude cuesta $200 al mes. Esto es GRATIS
Y viene con una idea brutal
Todo es un plugin
El modelo
Las tools
El sandbox
La UI
Hasta el propio loop del agente
Puedes cambiar cualquiera de esas piezas solo modificando la configuracion
Sin tocar el codigo base
Un solo comando y tienes la interfaz web local:
npx @deepseek-ai/dsh web
Compatible con DeepSeek, Claude, GPT, Gemini y lo que quieras enchufar
Llego a mas de 160 mil estrellas en pocos dias
Un laboratorio frontier acaba de soltar gratis la capa exacta que otras empresas estan vendiendo como producto premium
Te dejo el repo en los comentarios
Andrej Karpathy spent 8 years at OpenAI and Tesla
Last week, he put everything he knows into one free 2-hour lecture
People pay $15k for bootcamps that teach half of this
You probably don't have 2 hours right now
Don't let this get lost in your feed
Watch it, then read the guide below and build your first loop
An absolute banger of a paper.
"A Gentle Introduction to Matrix Calculus" by econometrics legend Jan Magnus — one of the clearest explanations of matrix derivatives ever written. Published in the Journal of Econometrics in 2024.
If you work in econometrics, machine learning, statistics or optimisation, this paper is pure gold.
Free, in my Awesome Math Books list (econometrics section)
https://t.co/soOYrEQ2he
I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.
https://t.co/NHyibLgade
Predicting the answer to interventional "what if?" questions — the outcome of an action you never took — need a *mechanistic* model, not a curve fit. And you can only learn one by *experimenting*. Experiments are costly, so the real game is **data efficiency**.
Meet the Model Discovery Agent (MDA). 🧵
Google just released a free 2-hour course on full Graph Engineering.
How to go from one prompt to 100 agents running inside one graph:
17:44 - Build your first AI agent
39:30 - Run agents with loop engineering
1:12:38 - Turn agent loops into graphs
1:34:26 - Build agents that throttle themselves
1:55:05 - Orchestrate the full multi-agent system
Most people build one agent and stop there.
Google is teaching the full stack:
Prompt → Agents → Loops → Graphs → Multi-Agent Systems
Single agents are the old workflow.
Self-regulating agent graphs are the new one.
This 2-hour watch is worth more than most paid agent engineering courses.
Bookmark and watch it today
Then read the full architecture below ↓
Modern AI is based on artificial neural nets (NNs). Who invented them? https://t.co/ZCI8ZrEKnZ
Biological neural nets were discovered in the 1880s [CAJ88-06]. The term "neuron" was coined in 1891 [CAJ06]. Many think that NNs were developed AFTER that. But that's not the case: the first "modern" NNs with 2 layers of units were invented over 2 centuries ago (1795-1805) by Legendre (1805) and Gauss (1795, unpublished) [STI81], when compute was many trillions of times more expensive than in 2025.
True, the terminology of artificial neural nets was introduced only much later in the 1900s. For example, certain non-learning NNs were discussed in 1943 [MC43]. Informal thoughts about a simple NN learning rule were published in 1948 [HEB48]. Evolutionary computation for NNs was mentioned in an unpublished 1948 report [TUR1]. Various concrete learning NNs were published in 1958 [R58], 1961 [R61][ST61-95], and 1962 [WID62].
However, while these NN papers of the mid 1900s are of historical interest, THEY HAVE ACTUALLY LESS TO DO WITH MODERN AI THAN THE MUCH OLDER ADAPTIVE NN by Gauss & Legendre, still heavily used today, the very foundation of all NNs, including the recent deeper NNs [DL25].
The Gauss-Legendre NN from over 2 centuries ago [NN25] has an input layer with several input units, and an output layer. For simplicity, let's assume the latter consists of a single output unit. Each input unit can hold a real-valued number and is connected to the output unit by a connection with a real-valued weight. The NN's output is the sum of the products of the inputs and their weights. Given a training set of input vectors and desired target values for each of them, the NN weights are adjusted such that the sum of the squared errors between the NN outputs and the corresponding targets is minimized [DLH]. Now the NN can be used to process previously unseen test data.
Of course, back then this was not called an NN, because people didn't even know about biological neurons yet - the first microscopic image of a nerve cell was created decades later by Valentin in 1836, and the term "neuron" was coined by Waldeyer in 1891 [CAJ06]. Instead, the technique was called the Method of Least Squares, also widely known in statistics as Linear Regression. But it is MATHEMATICALLY IDENTICAL to today's linear 2-layer NNs: SAME basic algorithm, SAME error function, SAME adaptive parameters/weights. Such simple NNs perform "shallow learning," as opposed to "deep learning" with many nonlinear layers [DL25]. In fact, many modern NN courses start by introducing this method, then move on to more complex, deeper NNs [DLH].
Even the applications of the early 1800s were similar to today's: learn to predict the next element of a sequence, given previous elements. THAT'S WHAT CHATGPT DOES! The first famous example of pattern recognition through an NN dates back over 200 years: the rediscovery of the dwarf planet Ceres in 1801 through Gauss, who collected noisy data points from previous astronomical observations, then used them to adjust the parameters of a predictor, which essentially learned to generalise from the training data to correctly predict the new location of Ceres. That's what made the young Gauss famous [DLH].
The old Gauss-Legendre NNs are still being used today in innumerable applications. What's the main difference to the NNs used in some of the impressive AI applications since the 2010s? The latter are typically much deeper and have many intermediate layers of learning "hidden" units. Who invented this? Short answer: Ivakhnenko & Lapa (1965) [DEEP1-2]. Others refined this [DLH]. See also: who invented deep learning [DL25]?
Some people still believe that modern NNs were somehow inspired by the biological brain. But that's simply not true: decades before biological nerve cells were discovered, plain engineering and mathematical problem solving already led to what's now called NNs. In fact, in the past 2 centuries, not so much has changed in AI research: as of 2025, NN progress is still mostly driven by engineering, not by neurophysiological insights. (Certain exceptions dating back many decades [CN25] confirm the rule.)
Footnote 1. In 1958, simple NNs in the style of Gauss & Legendre were combined with an output threshold function to obtain pattern classifiers called Perceptrons [R58][R61][DLH]. Astonishingly, the authors [R58][R61] seemed unaware of the much earlier NN (1795-1805) famously known in the field of statistics as "method of least squares" or "linear regression." Remarkably, today's most frequently used 2-layer NNs are those of Gauss & Legendre, not those of the 1940s [MC43] and 1950s [R58] (which were not even differentiable)!
SELECTED REFERENCES (many additional references in [NN25] - see link above):
[CAJ88] S. R. Cajal. Estructura de los centros nerviosos de las aves. Rev. Trim. Histol. Norm. Patol., 1 (1888), pp. 1-10.
[CAJ88b] S. R. Cajal. Sobre las fibras nerviosas de la capa molecular del cerebelo. Rev. Trim. Histol. Norm. Patol., 1 (1888), pp. 33-49.
[CAJ89] Conexión general de los elementos nerviosos. Med. Práct., 2 (1889), pp. 341-346.
[CAJ06] F. López-Muñoz, J. Boya b, C. Alamo (2006). Neuron theory, the cornerstone of neuroscience, on the centenary of the Nobel Prize award to Santiago Ramón y Cajal. Brain Research Bulletin, Volume 70, Issues 4–6, 16 October 2006, Pages 391-405.
[CN25] J. Schmidhuber (AI Blog, 2025). Who invented convolutional neural networks?
[DEEP1] Ivakhnenko, A. G. and Lapa, V. G. (1965). Cybernetic Predicting Devices. CCM Information Corporation. First working Deep Learners with many layers, learning internal representations.
[DEEP1a] Ivakhnenko, Alexey Grigorevich. The group method of data of handling; a rival of the method of stochastic approximation. Soviet Automatic Control 13 (1968): 43-55.
[DEEP2] Ivakhnenko, A. G. (1971). Polynomial theory of complex systems. IEEE Transactions on Systems, Man and Cybernetics, (4):364-378.
[DL25] J. Schmidhuber. Who invented deep learning? Technical Note IDSIA-16-25, IDSIA, November 2025.
[DLH] J. Schmidhuber. Annotated History of Modern AI and Deep Learning. Technical Report IDSIA-22-22, IDSIA, Lugano, Switzerland, 2022. Preprint arXiv:2212.11279.
[HEB48] J. Konorski (1948). Conditioned reflexes and neuron organization. Translation from the Polish manuscript under the author's supervision. Cambridge University Press, 1948. Konorski published the so-called "Hebb rule" before Hebb [HEB49].
[HEB49] D. O. Hebb. The Organization of Behavior. Wiley, New York, 1949. Konorski [HEB48] published the so-called "Hebb rule" before Hebb.
[MC43] W. S. McCulloch, W. Pitts. A Logical Calculus of Ideas Immanent in Nervous Activity. Bulletin of Mathematical Biophysics, Vol. 5, p. 115-133, 1943.
[NN25] J. Schmidhuber. Who invented artificial neural networks? Technical Note IDSIA-15-25, IDSIA, November 2025.
[R58] Rosenblatt, F. (1958). The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review, 65(6):386.
[R61] Joseph, R. D. (1961). Contributions to perceptron theory. PhD thesis, Cornell Univ.
[R62] Rosenblatt, F. (1962). Principles of Neurodynamics. Spartan, New York.
[ST61] K. Steinbuch. Die Lernmatrix. (The learning matrix.) Kybernetik, 1(1):36-45, 1961.
[TUR1] A. M. Turing. Intelligent Machinery. Unpublished Technical Report, 1948. In: Ince DC, editor. Collected works of AM Turing—Mechanical Intelligence. Elsevier Science Publishers, 1992.
[STI81] S. M. Stigler. Gauss and the Invention of Least Squares. Ann. Stat. 9(3):465-474, 1981.
[WID62] Widrow, B. and Hoff, M. (1962). Associative storage and retrieval of digital information in networks of adaptive neurons. Biological Prototypes and Synthetic Systems, 1:160, 1962.
Introducing "Diffusing Blame": can a neural network learn competitively while strictly obeying Dale's principle, the rule that real neurons follow? We show it can, across both image classification and reinforcement learning. 🧠
Accepted at #ALIFE2026
https://t.co/oSsfRCpvc7
Real neurons generally follow Dale’s principle: each neuron is predominantly excitatory or inhibitory. Standard artificial networks usually ignore this constraint, allowing every unit to mix positive and negative outgoing weights.
Backprop makes the gap even wider. Its backward pass needs exact transposed copies of the forward weights, the so-called "weight transport problem,” which biology doesn’t seem to have a mechanism for.
So we asked: can a network that strictly enforces Dale's principle still learn well, without weight transport?
Our approach builds on Error Diffusion (ED), a local rule that routes a single global error signal directly to every hidden unit, where each layer is split into separate excitatory and inhibitory streams with four non-negative weight matrices, so a synapse's sign comes from fixed population identity rather than a learnable weight. Our main contribution is to extend ED from binary to multi-class problems via modulo error routing.
We then asked whether this routing mechanism could provide useful credit signals in the noisy setting of reinforcement learning. During PPO training on Ant, Humanoid, and HalfCheetah, we compared each local ED update with the corresponding true backpropagation gradient. Among the routing schemes we tested, modulo routing consistently produced the strongest alignment.
Taken together, these results show that Dale-constrained networks can still learn without transporting weights backward, suggesting a potential path toward learning rules that are both effective and more biologically plausible.
new post on harness engineering for AI self-improvement: https://t.co/ZYvGfVs61k
It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple.
Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
"The role of raw power in intelligence", Hans Moravec, 1976 (!).
"The first section discusses natural intelligence, and notes two major branches of the animal kingdom in which it evolved independently, and several offshoots. The suggestion is that intelligence need not be so difficult to construct as is sometimes assumed.
The second part compares the information processing ability of present computers with intelligent nervous systems, and finds a factor of one million difference. This abyss is interpreted as a major distorting influence in current work, and a reason for disappointing progress.
Section three examines the development of electronics, and concludes the state of the art can provide more power than is now available, and that the gap could be closed in a decade."
https://t.co/NnoGCVh0LR
Andrew Ng just dropped a free course on Claude Code from scratch, taught with the Anthropic team:
00:00 - why Claude Code is so agentic
04:00 - shockingly simple architecture
12:00 - point it at any codebase
this short watch will replace 10 paid coding agent courses.
Andrew Ng calls it his personal favorite coding assistant right now.
Watch it today, then read how to engineer your own agent loops in the article below
VLMは人間のような創造性を持てるか?
ケネス・スタンレー教授らの『目標という幻想(Why Greatness Cannot Be Planned)』は、明確な目標を設定することが、かえって真に偉大な発見を遠ざけてしまうという逆説を論じた書籍です。その議論の中核にあったのが「PicBreeder」の実験でした。
PicBreeder では、ユーザーが「面白い」と感じた画像を選び、それを少しずつ進化させていきます。事前に決められたゴールはなく、人々が「なんとなく良い」と思ったものを選び続けるだけで、顔や動物、乗り物、頭蓋骨といった予期しない形が、何世代もかけて、多くの人の手を経て自然と現れます。スタンレー教授らは、こうした「オープンエンド」、つまり目標をあらかじめ定めない探索こそが人間の創造性の根幹にあると考えたのです。
では、このオープンエンドな探索を、AIは再現できるのでしょうか。
MIT・NYUとの共同研究として発表する「In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models」では、視覚言語モデル(VLM)エージェントによる再現を試みました。エージェントたちは共有アーカイブを探索し、画像を選んで進化させ、気に入ったものを公開し、他のエージェントの作品を評価します。目標となる画像も「進歩」の定義も与えられていません。
その結果、AIによるオープンエンドな発見の可能性と限界の両方が浮かび上がりました。VLMエージェントは特定の見た目や意味に引き寄せられやすく、既存のアイデアを捨てて予期せぬ何かを探すよりも、手元にあるものを洗練させることに留まりがちでした。一方で、多様な人格を持つエージェント集団を導入すると探索は大きく改善され、生成されたアーカイブの意味的な幅広さは、人間が作ったアーカイブに迫る水準にまで達しました。
しかし、VLMでは届かなかった点もありました。人間は、偶然の産物を持続的な創造へとつなげることに長けています。あるものを見つけたとき、その価値を感じ取り、それを追いかけることで、より大きな概念的飛躍を達成できます。AIエージェントも興味深いパターンに気づくことはできても、そのパターンに囚われてしまう傾向がみられました。
なぜ人間は、本研究のVLMにはできなかったオープンエンドな探索を進められるのか。現在のAIシステムに何が欠けているのか。私たちはまだ十分に理解できていません。ここには、当社が探求するAI駆動型科学研究にも通じる大きな問いが残っています。Sakana AIは今後も、オープンエンドな知性の探求を深めていきます。
ブログ:https://t.co/qsMwcB3N5D
論文:https://t.co/QnxVWLzjez 🐟
Are neural nets across modalities really converging to the same representation as they scale, as the Platonic Representation Hypothesis suggests?
We show that common representational similarity metrics are confounded by network width & depth. We propose a permutation-based null calibration that fixes this.
Result❓
• Global convergence largely disappears.
• Local neighborhoods persist.
We propose the alternative Aristotelian Representation Hypothesis: Neural networks, trained with different objectives on different data and modalities, are converging to shared local neighborhood relationships
Very proud of @FabianGroger and @ShuoWen18 for this work!
Paper: https://t.co/GmkhwsiN1N
Webpage: https://t.co/xaI31BU2FS
Code: https://t.co/5qItdzRBZP
J-space is really something we have been exploring since 2022. Glad to see it continues to work well at scale!
Some of the related work along this direction:
- How to recover the latent process using Jacobians (Identifiability of nonlinear ICA): https://t.co/uZkCxCMyrq
- How to handle dependent latents and assumption violations (again, through Jacobians): https://t.co/JNFPThjvkW
- For general latent variable models, what remains recoverable with guarantees, and why Jacobians are universally helpful? (We could generalize SAEs to the general nonlinear case, with Jacobians!): https://t.co/Mmh9lgglA0
Feels like we're still only beginning to uncover what Jacobian structure can tell us about representations.
🚀 We introduce Neural Theorizer (NEO) — a new type of world model that learns to theorize the world from observation, without language or LLM supervision.
Selected as an ICML 2026 oral presentation — 0.7% of submitted papers.
The paper asks:
"What does it mean to understand the world and build a world model?"
Today’s world models are often trained to predict the future: the next frame, next latent state, or next observation.
But is prediction enough?
We argue that a world model should be a theory-building system: one that discovers reusable primitives, composes them into executable explanations, and transfers those explanations to novel phenomena.
NEO is our first step toward this vision — a World Theory Model that learns explicit, compositional theories from raw observation.
This work was led by my wonderful students: Doojin Baek*(@doojin_a_baek), Gyubin Lee* (@gyubin0521), Junyeob Baek (@JunyeobB), and Hosung Lee (@HosungLee_).
For more details, take a look at the paper — and if you’re attending ICML, let’s talk there!
📄 arXiv: https://t.co/TGMXLLfzP7
🌐 Project page: https://t.co/aLJywp8rfq
Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API.
Our ‘Fugu Ultra’ model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls.
Try it: https://t.co/hhO6qTawgb 🐡
What if a robot policy weren't a neural net or a test-time chat loop, but a multi-file code repo selected from a Pareto frontier of genetically evolved candidates? RHO moves all its LLM exploration to training time, then runs that repo on scenes it was never trained on. 🧵👇🏽