Today we're introducing TRIBE v2 (Trimodal Brain Encoder), a foundation model trained to predict how the human brain responds to almost any sight or sound.
Building on our Algonauts 2025 award-winning architecture, TRIBE v2 draws on 500+ hours of fMRI recordings from 700+ people to create a digital twin of neural activity and enable zero-shot predictions for new subjects, languages, and tasks.
Try the demo and learn more here: https://t.co/VkMd1YpQWI
nanochat now trains GPT-2 capability model in just 2 hours on a single 8XH100 node (down from ~3 hours 1 month ago). Getting a lot closer to ~interactive! A bunch of tuning and features (fp8) went in but the biggest difference was a switch of the dataset from FineWeb-edu to NVIDIA ClimbMix (nice work NVIDIA!). I had tried Olmo, FineWeb, DCLM which all led to regressions, ClimbMix worked really well out of the box (to the point that I am slightly suspicious about about goodharting, though reading the paper it seems ~ok).
In other news, after trying a few approaches for how to set things up, I now have AI Agents iterating on nanochat automatically, so I'll just leave this running for a while, go relax a bit and enjoy the feeling of post-agi :). Visualized here as an example: 110 changes made over the last ~12 hours, bringing the validation loss so far from 0.862415 down to 0.858039 for a d12 model, at no cost to wall clock time. The agent works on a feature branch, tries out ideas, merges them when they work and iterates. Amusingly, over the last ~2 weeks I almost feel like I've iterated more on the "meta-setup" where I optimize and tune the agent flows even more than the nanochat repo directly.
Opus-4.5(T) found a specific pixel scaling pipeline issue in an entire VLM automation framework and fixed it to work with Qwen3-VLβs normalized pixel coordinates for output bounding boxes. This niche anomalyβs fix had been WIP for weeks. It one-shotted the fix.
Terrifyingly good
@Stagehanddev Great work team! Does v3 allow mapping individual aisdk based custom models to various actions (act, observe, extract)? More flexibility w.r.t open vision models apart from first class supported closed models would greatly improve DX.
Built @getbridgenow to solve a problem that is very close to my heart - fragmented application discovery and confusing interfaces.
Bridge solves it π
Experience an agentic hands-off experience, pay with your card/upi or stables on solana across apps.
https://t.co/878FlXW842
Introducing Bridge π
A sneak peek into the future of agentic experience between human intents and software actions.
Join the waitlist - https://t.co/nux8U3Np4i
Hold tight, we are bridging soon.
See ya all at @colosseum@solana@superteam
@shek_dev omniparser v2 (by microsoft) + UItars/gpt-4o runs inside a windows vm sandbox. Works fine with terminal, but a more generalized agentic computer use. https://t.co/NqXHiABX62
@kashdhanda@superteam Thanks for an amazing trek, Sherpa! You've shown countless talented people their first trails to the Solana peak. All the best for the next mountain to scale! β¨