We built an all-synthetic simulation of a hospital that can be shared publicly without any issues. Our data quality is so high, physicians cannot tell the difference between real/synthetic patients. The data is fully verified -- perfect for RL.
Introducing Synthetic Hospital: an open, fully synthetic longitudinal EHR benchmark with verifiable ground truth!
1,268 patients, 5,602 encounters, zero PHI. Physicians could not reliably distinguish its charts from real ones.
📄 https://t.co/kU6Rkqr2dA
💻 https://t.co/k3UH8vgQ6H
✍️ https://t.co/otOqD6LMee
CliffCompaction, our autocompaction tool for coding agents, is out!
Sessions run for millions of tokens, agents stay on track, and costs drop by up to 50%.
Paper: https://t.co/OdPp6Mt1Gh
Code: https://t.co/39Mg0ZJZ9C
PyPI: https://t.co/yOSIP3KCUH
Blog: https://t.co/cZdN9ZbURf
Our 2nd release in our open-source week.
CliffCompaction lets you run sessions for millions of tokens and save 50% on token costs. It integrates easily with Claude Code and Codex. We have been using it for months, and it has better vibes than Codex/Claude Code compaction.
Huge thanks to my advisors and coauthors @Tim_Dettmers, @EulrangCho, and Bingqing Chen. Thanks also to my labmates, collaborators, and friends who tried CliffCompaction and gave us feedback!
We’ve been using it in the lab for months, and I personally haven’t gone back to working without it. We hit usage limits much less often, and we’ve observed that agents forget less and stay more on track. Our partners have deployed it across their company and seen cost reductions in token spend of 45%.
https://t.co/1iTca1IILx
My first robotics paper is out 🤖
World models are basically robotics' smartest friend who thinks 7 times before answering any question. great with text, terrible at parties 🥺
So we asked: can we keep the brain and lose the lag? THAW-VLA 🧵
(1/n)
This is the first release in our open-source week: runtime dynamic compression and the private beta for bitsandbytes2.
bnb2 offers a new level of compression that brings large models to smaller devices.
The bnb2 private beta has started. We offer a limited number of people access to an early version of bnb2.
Sign up below.
Releasing our runtime dynamic compression framework that achieves 1.5-2.0 bit compression at high quality. This is integrated into the bitsandbytes2 library, which starts as a private beta today.
Paper: https://t.co/f4L4Nm8JwM
Private beta signup: https://t.co/q8PPAulmq3
Starting tomorrow, our lab will hold an open-source week: 2 software frameworks, 4 papers, all building a coherent ecosystem.
The theme: Frontier AI on Hardware You Own
Blog post: https://t.co/e55yWC9PGJ
SoTA results in:
Autocompaction
Autonomous Research
Model compression
Test-time scaling
Deep Research
And a healthcare RL environment giving you a new level of complexity to train healthcare agents.
The ecosystem that we will release is built to be as usable as possible. Agent sessions that run overnight and go on for tens of millions of tokens is made easy. Model compression of a model is automatic: you just get a good model that runs fast locally -- no expertise required. An autonomous research that works out of the box.
Our ecosystem enables a new level of work that can be done locally.
basketball AI (95% local AI + 5% GPT-6 Astra)
- detect ball and players
- track players
- re-identify players across plays
- OCR player numbers
- recognize player in possession
- detect court keypoints
- map player positions and trajectories
Astra is powerful, but expensive tool. use it when it actually makes a difference.
↓ deep dive
I asked Astra to recreate @LewisHamilton’s Ferrari!
It went through F1 specs, regulations, videos from F1 engineers and photos of the car, then built a 3D model (w/ blender) with a full breakdown of the parts.
It even made clay & wireframe versions too.
Every time I think I’ve pushed this model to its limits, it surprises me.