🚨BREAKING: Princeton just proved that AI agents are throwing away
the most valuable data they'll ever collect.
And nobody noticed because it looks like normal conversation.
Every time an AI agent takes an action, it receives what researchers
call a "next-state signal." A user reply. A tool result. A terminal
output. A test verdict.
Every existing system takes that signal and uses it as context for
the next response.
Then discards it forever.
The Princeton team just proved this is one of the most expensive
mistakes in AI engineering. Because that signal contains two things
nobody was extracting.
First: an implicit score. A user who re-asks a question is telling
you the agent failed. A passing test is telling you it succeeded.
A detailed error trace is scoring every step that led to it. This
is a live, continuous reward signal hiding inside every interaction.
Free. Universal. Completely ignored.
Second: a correction direction. When a user writes "you should have
checked the file first," they're not just saying the response was
wrong. They're specifying which tokens should have been different
and how. That's not a scalar reward. That's token-level supervision.
And scalar rewards throw every single bit of it away.
They built a system called OpenClaw-RL around recovering both.
Then they ran the experiment that changes everything.
An agent started with a personalization score of 0.17. After just
36 normal conversations, with no new training data, no labeled
dataset, and no human annotations, the combined method hit 0.81.
The agent didn't get retrained. It got used.
That's the part nobody is talking about. The model was serving live
requests at the same time it was being trained on them. Four
completely decoupled loops running simultaneously. Policy serving.
Rollout collection. Reward judging. Weight updates. None waiting
for the others.
The agent gets smarter every time someone talks to it.
And the deeper the task, the more it matters. On long-horizon
agentic tasks, outcome-only rewards give you a signal at the very
end of a trajectory and nothing in between. Their process reward
model scores every single step using the live next-state signal as
evidence. Tool-call accuracy jumped from 0.17 to 0.30. GUI accuracy
improved further on top of that.
This creates a shift nobody has fully reckoned with yet.
The current paradigm: collect data offline, train in batches,
deploy, hope it works.
The new paradigm: deploy, extract training signal from every
interaction, update continuously, improve automatically.
Every conversation is training data. Every correction is a gradient.
Every re-query is a reward signal.
The agents that figure this out first won't need bigger datasets.
They'll just need more users.
Can a single agent with skills replace multi-agent systems?
Multi-agent systems work well for complex reasoning where specialized agents collaborate through explicit communication.
But this incurs substantial computational overhead in tokens and latency.
This new research explores whether you can compile a multi-agent system into an equivalent single-agent system by trading inter-agent communication for skill selection.
The answer: yes, but with a caveat.
Preliminary experiments show single-agent approaches with skill libraries can substantially reduce token usage and latency while maintaining competitive accuracy on reasoning benchmarks.
So far, so good.
But here's where it gets interesting. The researchers asked: How does skill selection scale as libraries grow?
Drawing on cognitive science, they propose that LLM skill selection exhibits bounded capacity analogous to human decision-making. And they found an interesting pattern.
Rather than degrading gradually, selection accuracy remains stable up to a critical library size, then drops sharply. But it looks like a phase transition, not a smooth decline. This mirrors capacity limits observed in human cognition.
The culprit isn't library size alone. It's the semantic confusability among similar skills. When skills are too semantically similar, the model can't reliably distinguish between them.
This suggests hierarchical organization, which has long helped humans manage complex choices, may similarly benefit AI systems. Initial results with hierarchical routing support this hypothesis.
As we build increasingly capable agents with expanding skill sets, understanding these fundamental limits becomes critical. You can't just keep adding skills indefinitely. There's a threshold where selection breaks down, and it happens suddenly, not gradually.
Paper: https://t.co/3RMAu3Fcnp
Learn to build effective AI agents in our academy: https://t.co/zQXQt0PMbG
We just open sourced the code-simplifier agent we use on the Claude Code team.
Try it: claude plugin install code-simplifier
Or from within a session:
/plugin marketplace update claude-plugins-official
/plugin install code-simplifier
Ask Claude to use the code simplifier agent at the end of a long coding session, or to clean up complex PRs. Let us know what you think!
The paper says the best way to manage AI context is to treat everything like a file system.
Today, a model's knowledge sits in separate prompts, databases, tools, and logs, so context engineering pulls this into a coherent system.
The paper proposes an agentic file system where every memory, tool, external source, and human note appears as a file in a shared space.
A persistent context repository separates raw history, long term memory, and short lived scratchpads, so the model's prompt holds only the slice needed right now.
Every access and transformation is logged with timestamps and provenance, giving a trail for how information, tools, and human feedback shaped an answer.
Because large language models see only limited context each call and forget past ones, the architecture adds a constructor to shrink context, an updater to swap pieces, and an evaluator to check answers and update memory.
All of this is implemented in the AIGNE framework, where agents remember past conversations and call services like GitHub through the same file style interface, turning scattered prompts into a reusable context layer.
----
Paper Link – arxiv. org/abs/2512.05470
Paper Title: "Everything is Context: Agentic File System Abstraction for Context Engineering"
MIT asked 153 execs with enormous AI budgets what they actually want from an AI partner.
They didn’t want "cutting-edge AI" for the sake of it - nor some "revolutionary platform."
They just asked for someone who:
a) understands their actual workflows
b) can guide effective implementation
Literally internalize this chart:
This could be an incredible revolution in Cosmology.
The Dark Energy model of the universe, which won a Nobel Prize in 2011, may be completely wrong.
The accelerating expansion instead is simply because time runs faster in the voids between galaxies.
Let me explain:
Smol VLMs ftw! Microsoft just dropped Florence - SOTA 200M & 800M parameter vision foundation model! 🔥
> Best part MIT Licensed! 🤯
> 200M checkpoint beats Flamingo 80B (400x bigger model) by a huge margin
> Performs captioning, object detection and segmentation, OCR, phrase grounding and more
> Leverages FLD-5B dataset - 5.4 billion annotations across 126 million images
> Multi task learning
> Finetuned model checkpoints beat the likes of PaLI, PaLI-X
Thanks and kudos to Microsoft for choosing open source! 🤗
This paper can be MASSIVE to solve the KV cache memory problem, achieves up to 26�� higher throughput than standard transformers and competitive performance in language modeling and downstream tasks. 🔥
The KV cache can take up over 30% of the GPU memory. The paper proposes a solution to dramatically reduce the KV cache size and improve inference throughput, by caching the KVs of just one layer (the top layer), by exploiting the fact that the topmost layer's representation is most informative 💯
"Layer-Condensed KV Cache for Efficient Inference of Large Language Models" ✨
📌 The core idea is to only compute and cache the keys and values of a small number of layers, instead of all the layers. Specifically, the queries of all layers attend to the keys and values of only the topmost layer.
📌 By caching the KVs of just one layer (the top layer) instead of all layers, the memory consumption is significantly reduced. Furthermore, since the keys and values of the other layers are no longer needed, their computation can be skipped entirely and their corresponding weight matrices (W_K and W_V) can be discarded. This not only saves more memory but also improves throughput by avoiding unnecessary computations.
📌 To prevent a drop in model performance, a small number of "warmup" layers are retained that use standard attention. These warmup layers are placed in a "sandwich" style - half at the very bottom of the model and half at the very top. Experiments show this sandwich placement works best.
📌 Training this modified architecture is tricky because the computation of each token now depends on the top-layer KVs of previous tokens, preventing parallelization. The authors derive an approximate parallel training method where KVs are iteratively refined over multiple rounds, with backpropagation only through the last 2 rounds.
📌 Empirically, the KVs converge very quickly over these iterative refinement rounds, so only a small number of rounds (e.g. 7) are needed in practice before the final backpropagation rounds. More warmup layers lead to even faster KV convergence.
---
Two issues with this technique, that were noted in paper-
1. Longer training time - which may be mitigated by using a method similar to the one in the recent paper that tuned a pretrained Transformer into Feedback Transformer.
2. When prompt is long, inference performance degrades (could even be worse than before)
Today, we announced that we’ve gotten dictionary learning working on Sonnet, extracting millions of features from one of the best models in the world.
This is the first time this has been successfully done on a frontier model.
I wanted to share some highlights 🧵
From Claude100K to Gemini10M, we are in the era of long context language models. Why and how a language model can utilize information at any input locations within long context? We discover retrieval heads, a special type of attention head responsible for long-context factuality
Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding
paper page: https://t.co/jsAFklEnC9
Table-based reasoning with large language models (LLMs) is a promising direction to tackle many table understanding tasks, such as table-based question answering and fact verification. Compared with generic reasoning, table-based reasoning requires the extraction of underlying semantics from both free-form questions and semi-structured tabular data. Chain-of-Thought and its similar approaches incorporate the reasoning chain in the form of textual context, but it is still an open question how to effectively leverage tabular data in the reasoning chain. We propose the Chain-of-Table framework, where tabular data is explicitly used in the reasoning chain as a proxy for intermediate thoughts. Specifically, we guide LLMs using in-context learning to iteratively generate operations and update the table to represent a tabular reasoning chain. LLMs can therefore dynamically plan the next operation based on the results of the previous ones. This continuous evolution of the table forms a chain, showing the reasoning process for a given tabular problem. The chain carries structured information of the intermediate results, enabling more accurate and reliable predictions. Chain-of-Table achieves new state-of-the-art performance on WikiTQ, FeTaQA, and TabFact benchmarks across multiple LLM choices