What if,
for any scientific claim,
you can search 100M+ papers to find the strongest evidence for and against it
⚡️in 1 second?
That’s now possible with Evidence API! 🚀
https://t.co/2I25q4m9ao
For example, does drinking coffee reduce diabetes risk?
I propose Stanford NLP as an independent third-party evaluator under @DarioAmodei’s 3 step plan. For important parts of the work, universities would be better than any other organization (see below 🧵👇), and, of university groups, @stanfordnlp would be the best one to choose. 😊
Will be presenting our work on building a Universal Vision-Language World Model.
We build an 8B autoregressive transformer that unifies World Models and Vision-Language Models using visual abstractions as streams of thought, e.g., camera pose, depth, optical flow, point tracks, text -- generated and read like tokens. Our model solves vision-language tasks using world modeling (novel view synthesis) and inverse dynamics (camera pose estimation), unlike standard VLMs. We optimize end-to-end with a continuous rollout using RL, entering a cycle of recursive self-improvement.
New UI for controlling video generation, with more physical consistency. Pixels, along with visual abstractions like flow, depth, camera pose, become tokens to be conditioned on or generated.
We're releasing PSI-0.5: a promptable physical world model. Many world models let you explore by moving around a scene, PSI also lets you interact with it!
Excited to present our ICLR 2026 paper tomorrow: Unified 3D Scene Understanding Through Physical World Modeling.
Joint work with @KlemenKotar@Rahul_Venkatesh@jwhooglee@honglin_c@khai_loong_aw@dyamins.
We introduce 3WM, a foundation model for 3D understanding that treats depth, novel view synthesis, object motion, and geometric reasoning as different prompts to the same physical world model.
If you are at ICLR, come by our poster: Poster Session 6, Pavilion 4, Sat 3:15 PM.
Today's best AI needs orders of magnitude more data than a human child to achieve visual competence.
We introduce the Zero-shot World Model (ZWM), an approach that substantially narrows this gap. Even when trained on the first-person experience of a single child, BabyZWM matches state-of-the-art models on diverse visual-cognitive tasks – with no task-specific training, i.e., zero-shot. 🧵