Reka has signed the Open Weights and American AI Leadership letter.
We believe an open AI ecosystem is essential to advancing research, expanding access, and building a more innovative and resilient field. That's why we've openly released datasets like CS2-10k and Research-Eval, models like Reka Flash 3 and Reka Edge, and frameworks like Reka Quant.
https://t.co/7WDzElEvXZ
Today we're opening the WorldModelGym leaderboard 📊
Most world model evals ask one question: does the generated video look real? WMGym asks a different one. Give a model a menu of possible actions, can it correctly predict which one actually leads to the best outcome?🎯
We ran our own Dreamer-v3 across all four benchmark families as the first entry. It currently leads three of them — Meta-World, DeepMind Control, and Classical Control (where it hits 84% decision fidelity), and sits second on Atari.
We invite you to also enter your own model: you host a /score endpoint, we call it and compute the score ourselves.
Read our full breakdown and methodology: https://t.co/oqVMqJcHUT
My first interview with @sama, Co-Founder of @OpenAI.
0:04 How to start a startup
3:30 Trusting exponentials
4:57 Operating in chaotic environments
6:12 Learning to enjoy painful experiences
8:03 Creating abundant intelligence
11:15 Keeping core suppliers on OpenAI’s timelines
12:10 Invention of the joint-stock company
15:30 The best CEOs aren’t sociopaths
16:46 We are in the singularity
18:09 AI authoritarianism vs liberty
19:24 Texting 300-400 people a day
20:32 Having a small number of deep beliefs about the future
21:41 Critical path
22:25 Thinking about what’s next
23:38 Getting on planes in marginal situations
28:16 Buying lots of compute
30:51 Ambition
34:00 Google shouldn’t have let OpenAI survive
37:54 Having his life shot through a cannon after the launch of ChatGPT
41:20 The growth of Codex
42:02 The Death Star tweet
44:46 Status games and desire to be useful
47:57 Not being ambitious enough on compute investments
50:26 Execution
51:45 Ask for what you want
54:11 First few weeks of OpenAI
55:15 Shutting down Sora to focus on Codex
57:44 Designing beautiful products
59:01 Getting addicted to TikTok
1:01:03 Inventing a new device
1:02:53 Try to get better at your strengths
1:05:57 Masa is an n of 1
1:07:11 Real trends vs fake trends
Before a single model weight is updated, someone has to answer a deceptively hard question: What's actually in your data? 🤔
Julian Geoffrey López and Fedor Zhdanov share how they think about composition, quality, and annotation, and why getting this right determines what your world model knows, and what it doesn't.
🎥 Watch the full version: https://t.co/h31Uclx08X
📄 Read more: https://t.co/BBily52xyx
#PhysicalAI #WorldModels #DataPipeline #AI #RekaLabs
@KonradJam leads the data platform team at Reka. His job: prepare hundreds of thousands of hours of video data for world model training, keeping pace with a research team that moves fast.
In this episode, he and Julian Lopez share how the data pipeline for world model training actually works.
Less than 100 people. Petabytes of data. No two days are the same.
📄Learn more about the work we do → https://t.co/9SE1HnO0mD
#PhysicalAI #WorldModels #DataPipeline #AI #RekaLabs
🤔What does it take to train an omni world model from scratch?
📽️Petabytes of video. 6⃣pipeline stages. And every improvement in data quality pays off twice when you're training a model that both generates and understands video.
Our Reka Labs' data team on how it works → https://t.co/BBily52xyx
#PhysicalAI #WorldModels #AI #RekaLabs
In this video, @MateuszOnAI and @zaheri_h_r sat down to talk through two questions driving recent Reka Labs research.
1️⃣ Can VLMs understand physics the way we do?
2️⃣ And if an agent uses a world model to choose between actions, does it pick the right one?
Two benchmarks, two complementary questions about what it actually means for a model to understand the physical world.
Read more on the benchmarks here:
📄 PhysicalRealismBench-U → https://t.co/z6pPl2ly8S
📄 WorldModelGym → https://t.co/9n2ei3LuyY
new post on harness engineering for AI self-improvement: https://t.co/ZYvGfVs61k
It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple.
Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
Most evaluations of world models ask, "Does this look right?" 🌎
We built WorldModelGym to ask a different question: if an agent actually uses this model to make decisions, does it still choose well? 🏋️🏃♂️🏁🦾
A simulation can look physically plausible and still lead an agent completely astray. 📉⚠️
Stay tuned. More tomorrow at https://t.co/gaf6iEgEiw. 👀
🎮🕹️🖥️ CS2-10k is now available on @huggingface 🚀
600,000+ egocentric gameplay videos. 10,000+ hours.
Every frame paired with the exact keyboard, mouse, and 3D position data that produced it.
If you're working on world models, action-conditioned video generation, or egocentric navigation, this is ready to download and use today.
A super long overdue (3+ years?) post on scaling laws.
Compute is expensive. Scaling laws are a way to help us reason about the optimal compute allocation between data and model size before committing to a large run.
The post covers what scaling laws predict, how compute-optimal allocation works, why Kaplan et al. and Chinchilla disagree, and how data limits + fitting details make extrapolation tricky.
https://t.co/HP26eJvjHB
Training world models needs egocentric video and dense action signals, synchronized. That data is genuinely hard to find.
We built it from Counter-Strike 2 demos.
CS2-10k: 600K+ player-round videos, 10K+ hours, per-frame annotations — keyboard state, mouse delta, 3D position, camera yaw/pitch. All paired to the visual stream.
Why CS2 demos? Matches are deterministic replays. We can reconstruct first-person video and extract the exact control inputs that caused every visual change. No labeling, no estimation.
We're also releasing cs2-dem-renderer — the open-source pipeline we used to build it. Give it a .dem file, it outputs .mp4 + .parquet. Run your own dataset at whatever scale you need.
We taught a brand-new mini-series this year at @SCSatCMU on Modern GPU Programming for ML Systems, as part of the ML Systems course, touching on fun questions like what data layout swizzling is, how to use 3D TMA, and state-of-the-art Blackwell programming. We released a curated online book based on the materials: https://t.co/5ZJg2lySNO check it out
Introducing LifeSciBench, a benchmark for measuring and improving how well AI supports real-world life science research.
Developed with 173 scientists from biotechnology and pharmaceutical research, LifeSciBench includes 750 expert-authored tasks across seven biological research workflows.
https://t.co/JTk0wXHFrT
Let’s talk about evals.
We’re always looking for better ways to measure and forecast model progress, especially as benchmarks get saturated or gamed.
@tejalpatwardhan, who leads our frontier evals team, spoke to @andrewmayne about why evals matter and what models need to be judged on next.
We just released PhysicalRealismBench-U — a benchmark for testing whether VLMs actually understand physics in programmatically generated videos, fully attributable.
This is an important step toward models that understand and generate physically realistic outputs.
Best result across 9 frontier models: 57.7% realism F1.
Read our blog post here: https://t.co/yOEyZ2CASA
Visit the benchmark: https://t.co/Z9gPr4CiYt