COLM 2026 Main Accepted🎉!
Is AI good at predicting how groups will behave in the future?
We know AI can make such predictions. But how well can it actually do? And more importantly, how can we make it better?
In our new work Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis we study this problem systematically. Paper link: https://t.co/2k87ulCkx7
If you're interested in AI for future prediction—especially predicting the behavior of companies, organizations, governments, and other organized groups—check out our work:
🔗 https://t.co/zFFXgJkKms
Huge thanks to my amazing collaborators @yeeelow233 , @JasonWuzh@NanHuang99 and other collaborators, and to my advisors @LetianPeng and @shangjingbo for their invaluable guidance
🚀 Reef now supports training with the Tinker API!
Unlock continual learning with LoRA training by changing just one line in your config. Collect feedback, train, and publish versioned updates. No local GPU required.
Thanks to @Jayzou3773 for implementing this feature! 🙌
Check out our repo and contributions are always welcome! 👇
https://t.co/IOhoLExXmI
(1/6) Open Weight FastVideo FastH3 V2! Up to 9x speedup on @NVIDIA Blackwell. Lossless quality! 🚀
Lossless quality compared to base 50-step @MiniMax_AI H3. Judge for yourself below!
- FastVideo collab w/@nuvalab, NVIDIA FastGen, NVIDIA Enterprise Products
- Day 0 @ComfyUI weights and workflows
- Day 0 API served on @reactorworld
- Omni-ref currently training 👨🍳
More comparisons below!
We are excited to share ssa: a social simulation arena for evaluating your agent in the real future simulation.
Come and check it out:
https://t.co/ywi6yusOPj
Social simulation is now a product. Evaluation still has no common standard.
What if we tested every simulator against the real future?
We built Social Simulation Arena with researchers from MIT, Stanford, Harvard, CMU, UC Berkeley, and beyond.
Persona-prompted populations. Digital twins. Silicon samples. Synthetic populations. Worlds of a billion agents.
We have more ways than ever to simulate human behavior. But we still lack a standard everyone can trust to tell which ones work.
Most studies use their own datasets, often historical surveys whose answers are already online. A model that has seen the answer sheet can look brilliant.
Social Simulation Arena is a different kind of test: prospective, independent, and shared.
Simulators submit and lock their forecasts before each data release. When the real results arrive, every entrant is scored by the same rules.
Can a simulator predict a public before it moves?
If you’re building one, bring it.
https://t.co/SIytJoXeIR
#MIT #Stanford #Harvard #AI #Agents #LLM #SocialSimulation #Evaluation
In 1852 aluminium cost twice as much as gold.
Then the price fell over 99.9%. Yet it built one of the largest materials markets on earth.
What it says about AI tokens: https://t.co/5Y04wrz8mj
Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet?
Meet TRACES 🧭 — the world's first benchmark for measuring discoverative AI: AI that can work through evidence, test hypotheses, and reach verifiable conclusions on problems without answer keys. Proposed by our founder @tianqiao_chen, who defined its six capabilities.
Three things published today: a definition of "discoverative intelligence", a rubric to tell sound investigation from lucky guesses, and a open call for both solvers and problems
*Website: https://t.co/iJ5rDV6qz9
COLM 2026 Main Accepted🎉!
Is AI good at predicting how groups will behave in the future?
We know AI can make such predictions. But how well can it actually do? And more importantly, how can we make it better?
In our new work Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis we study this problem systematically. Paper link: https://t.co/2k87ulCkx7
If you're interested in AI for future prediction—especially predicting the behavior of companies, organizations, governments, and other organized groups—check out our work:
🔗 https://t.co/zFFXgJkKms
Huge thanks to my amazing collaborators @yeeelow233 , @JasonWuzh@NanHuang99 and other collaborators, and to my advisors @LetianPeng and @shangjingbo for their invaluable guidance
Completion before perfection.
1. Every “perfect” thing is built on countless small completions.
2. Everyone has a different definition of a great project. A finished version is already good enough for most people.
3. In the AI era, there are simply too many things to build. The difference between “perfect” and “done” matters less and less.
The world moves too fast. Many opportunities won’t wait for perfection.