Shrimp Welfare Bench: I've tried growing pothos in the tank before (it's a great ammonia/nitrate sponge) but I took it out because I was worried about night oxygen levels. Now I have my new air stone I'm having another go. Hopefully it'll cut the algae.
Reports From The Future tracks and clusters what the AI world has been yapping about. This week:
@latentspacepod admires Gemini 4 Argon: https://t.co/MzrpbSuNx0
The @FTC investigates biglab crimes (kind of): https://t.co/oBkEE3Wt7P
@_yusufknl elucidates grokking: https://t.co/h9j09NMgyE
An @MIT team investigates the evolution of specialised cognitive architecture in models: https://t.co/bI4jkYLojG
See these and more here: https://t.co/xOLs7SDMBV
As someone who ships LLM systems in production, this grokking video is the closest thing to a "the OpenAI vacation accident that changed how we think about AI learning" explainer I've ever seen released for free.
Everyone thinks a model that stops improving is done learning. It isn't. An OpenAI researcher left a model training over vacation, and after 7,000 flat steps it suddenly hit the jackpot - jumping from pure memorization to perfect understanding.
Bookmark this 35-min video and watch tonight. Same "grokking" effect discovered by pure luck in 2021, now the reason no lab trusts a flat loss curve anymore.
For those who missed it yesterday, I wrote an explainer on diffusion policies in robot learning, as I'd have wanted them taught to me. It covers:
- Why VAEs can still blend demos, and how hierarchical VAEs can help
- Diffusion as a special case of a hierarchical VAE, and why it works
- How Diffusion Policy brings this to policy learning, and my results reimplementing and tweaking it (video)
- How DDIM takes fewer steps, and how flow matching straightens its path into a simpler recipe (and reframes diffusion)
Now on to sim-to-real RL with a physical SO-101. Will be sharing that here as I go.
Can complex multi-step reasoning emerge purely from cells that only talk to immediate neighbors?
Happy to share our paper “Reasoning with Neural Cellular Automata (NCA)”, from our team at Google, Paradigms of Intelligence 🧵👇
Trying to gauge interest for a rugged computer in a box. We designed this for the mining industry but if things get much worse maybe there's a wider market for bugout compute.
This one has an RTX 4000 Blackwell in. Real prototype left, projected product right.
Every week we pull out and analyse the AI topics that have been grabbing mindshare over the past seven days. This week it picked up on:
@icarolab_ai explores the geometry of undefined terms: https://t.co/5k54joyBu2
Datasette adds telemetry options: https://t.co/U1B5RTrZKP
@solidSF tinkers with Jev alternatives: https://t.co/LtAspkHmeb
A world-modelling team from HK picks an unfortunate name: https://t.co/rzYEeheZkr
@sebuzdugan eavesdrops on a conversation between models: https://t.co/vdt4oQyDG0
4/9
How do you study a distinction you cannot think?
Find its recurring structure. Map its geometry. Intervene.
Steer it, patch it, remove it. Test what changes across unfamiliar contexts.
We can begin investigating before we know how to name what we have found.
If you're not OpenAI or Anthropic, you're benchmarking to work out which model/harness is best for your tasks, not to show off. You need dynamic benchmarking. https://t.co/6FtsDH5Ufi
What mineral exploration processes can be automated or assisted with AI? We're trying to create as comprehensive an index as possible. If we missed your workflow or your tools, let us know. https://t.co/b2S4s32K7E