Introducing ApprenticeBench: computer use + continual learning on a real job.
We show Fable 5.1 and GPT-6 Astra can now continually learn on a job and surpass human professionals. A decisive step change in AI's job readiness.
No FDEs. Agents deploy themselves into the job. 🧵
ApprenticeBench is probably one of the most differentiating benchmarks right now.
and it shows how large the gap between open and closed models still is on a real knowledge job:
Fable 5.1: 72%, $18.23 per task
Kimi K3: 18%, $25.83 per task
Fable 5.1 and Astra can now learn a knowledge job end to end.
It’s the culmination of rapid and comprehensive capability advances across computer use, continual learning with memory notes, and long-horizon agentic behavior over the past year. Together, these advances lead to a step change in AI’s job readiness.
It’s a bittersweet moment for me personally. I’ve been working on computer-use agents since 2022, back in the early days. I’ve always believed that the best way for agents to become truly helpful to us is for them to operate and learn the same way we do: use the same software through GUIs, and go on a lifelong learning journey on the job.
That’s the path of least resistance for AI adoption across industries. And that’s what drove us to found @NeoCognition.
What surprises me is how fast frontier AI is getting there. The rapid ramp in the CUA capability across generations of Claude and GPT models is a loud and clear sign. Fable 5.1 and Astra can now use GUIs just as well as APIs. Computer use is not a fundamental blocker anymore. We are also seeing early signs of the emergence of continual learning, even with a simple approach based on memory notes.
But there’s still a lot of work to do. Researchers need to understand and improve how agents learn, in both effectiveness and efficiency. Builders need to turn those capabilities into systems people can trust. Organizations need to bring workers into the process and ensure they share in the gains.
AI has crossed an important threshold in job readiness. It is inevitable that they will start doing real work across industries. But together, we have the opportunity to make that a benevolent future that benefits everyone. It is a future worth building.
Action space for agents gets redefined today!
We transition from few agentic steps to now automating the full job functions! Fascinating times!
Congrats @ysu_nlp and the @NeoCognition team!
Introducing ApprenticeBench: computer use + continual learning on a real job.
We show Fable 5.1 and GPT-6 Astra can now continually learn on a job and surpass human professionals. A decisive step change in AI's job readiness.
No FDEs. Agents deploy themselves into the job. 🧵
"Build a product that's native and trustless, then Bitcoiners will come."
@dntse explains why Babylon extends its native Bitcoin product from staking to collateral.
Babylon founder David Tse (@dntse) at SBC '26, Stanford:
"GOAT is another project which uses BABE for the bridging to L2."
BABE is the witness encryption scheme for Groth16 verification on Bitcoin, published in January by researchers at @UCBerkeley, @Stanford, @babylonlabs_io and @ByzntinResrch.
GOAT BitVM3 builds on it. Links below.
The future of bio is powered by faster data
Introducing the Medra AI Experimentalist: an agent that turns goals into experimental designs, learns from every result, and develops the next assay
Excited to collaborate with @DARPA and @NVIDIAHealth on the future of science
We trained a ~frontier Deep Research Agent on academic budget
> 32 H100s
> 8K synthetic samples
> fully open training infra + recipe (SFT, mid-training, RL)
> models of diff sizes (2B -> 35B) ready to use out of the box
This is yet another demonstration of how the frontier of AI is changing. We have reached a point where open models + a small capable team + a few hundred Ks can produce specialized models with ~frontier capabilities. The future of AI doesn’t have to be held in a chokehold by a handful of closed models.
We've open-sourced everything we've built and learned from this project. Hope it helps the community build more!
📌 Project: https://t.co/ZTsYE4JPqF
📌 Paper: https://t.co/UW0A9i4jA3
📌 Code: https://t.co/Bxyf12Q4j6
📌 Model Weights and Data: https://t.co/W15CLnEO1T
📌 Demo: https://t.co/pEx2FOqGZd
Amazing effort led by @jianxie_ (our 1st year student!!), Tianhe Lin, Zilu Wang. joint with @hhsun1 and the @osunlp team. thanks @amazon Xiangjun Wang for a gift that covers the compute and fruitful discussion.
The key missing piece for this positive-sum future is AI agents still don't have the continual learning capability needed to sustain this human-AI learning loop.
that's what @NeoCognition is trying to solve: agents that can indefinitely learn to specialize for any profession, organization, and individual.
A refreshing take on world models from @drfeifei
Grounding the taxonomy (renderer, simulator, planner) in the classic agent loop is brilliant imo. Also agree on the unified world model vision.
A few important questions still remain quite open:
1. The current discourse on world models is centered around the physical world. How to model the digital world?
We need digital world models as well, but they probably will take quite different forms (e.g., less on rendering, symbolic representation / language takes more prominence)
2. What level of fidelity is needed for simulation?
Most of the time, humans are probably not simulating the world (physical or digital) at a high fidelity. our simulation can be quite coarse-grained, lossy, and opportunistic. and that seems sufficient for most human activities.
3. Do world models need language? Is language only for communication?
Modern human civilization builds on symbolic representations like language. Language is as much a tool for representation as it is a tool for communication. If world models are about representing the world around us, should we abandon the most powerful representational tool humanity has ever created and try to learn from scratch?
Thrilled to celebrate my three amazing students officially graduating today! They’ve made incredible contributions to diffusion/world models, AI for math, and learning theory/statistics.
Zihan Ding (@Hanry65960814) -> ByteDance
Shange Tang (@sangertang1999) -> OpenAI
Jiawei Ge (@EmilyJge) -> Berkeley postdoc, then Cornell faculty
Can’t believe five years have flown by. Working with you all has been such a gift!
This is insane level of performance results - outperforming the next best by 10X!
Congrats on the ICML spotlight on Walrus!
Amazing @cosmo_shirley , @mikemccabe210 and entire @PolymathicAI team!
1/ Today with my colleagues @PolymathicAI, I'm excited to release our latest project, Walrus, a cross-domain foundation model for physical dynamics, into the world.
https://t.co/ihv1MZGQM3
Paper: https://t.co/d6ah9LO4ud
Git: https://t.co/s3p8qGhZQR
HF: https://t.co/RufaBD9eJk
Surreal to see this. These life-size displays sitting today at an exhibition in Venice were long ago small images on my screen during a single session of Picbreeder where I serendipitously evolved a car. Crazy journey for these images!
@michellearning is a force of nature. She and team @medra_ai are working on one of the coolest thing in robotics.
Proud to be one of the earliest supporter