The #1 thing I want from any model right now: speed
Intelligence is already good enough. But waiting 1-5 min per task is the worst possible window: too short for deep work, too long to just stare at the screen. So I end up scrolling X
Agents are making all of us more ADHD
The #1 thing I want from any model right now: speed
Intelligence is already good enough. But waiting 1-5 min per task is the worst possible window: too short for deep work, too long to just stare at the screen. So I end up scrolling X
Agents are making all of us more ADHD
We ran a head-to-head benchmark of 7 AI video systems: Tap8, Fable 5, Seedance 2.0, Google Veo 3.1, and others.
The clips in this video are pulled directly from that benchmark so you can see the difference yourself.
Hey guys! It’s William, I will be on X from now on.
I’ll be sharing what we’re building, what we’re learning, and some of our latest progress across Diffusion LMs, RSI, 3D generation, and cool human–machine interaction.
Excited to meet more people here!
Our bet: 1. codes-based video rather than pixel so it’s always accurate & naively interactive.
2. real-time rendering via ultra-fast, post-trained, Diffusion-LM (~1000+ token/sec).
3. vision-based RSI so it’s SOTA on quality (btw our dlm BEATS FABLE 5 on code-based video quality, check out our benchmarks ⬇️)
Today we're announcing Tap8: interactive video as the interface between humans and AI.
The model answers with a small runnable world instead of prose. Every element is exact, and every element is a handle: pull on one and the explanation reshapes around it. Facts render as live code, not hallucinated pixels: tokens for what is true, pixel diffusion for what it looks like.
Research bets: visual quality trained by recursive self-improvement, and real-time generation on a Diffusion-LM.
Preview coming soon. https://t.co/plppNgWHg2
The chatbox had a good run.
If text is too slow, why not pure video? Because pixels lie.
A model trained to reconstruct appearance learns what a whiteboard looks like, never whether the math on it is right. Its loss was never a function of the math. More scale buys a more convincing whiteboard, not a more correct one.
Correctness and appearance are different problems. Treat them as one and you get beautiful things that are wrong.
OH MY GOD 🤯
SIGMAZ AI JUST SHIPPED REAL-TIME INTERACTIVE VIDEO AT 1/20TH THE INFERENCE COST OF DIT.
Their eval infra hits 69.5% agreement with human visual-quality preference. Single-score review baselines sit at 48.5%.
Peer-reviewed at ICLR 2026, CTO Derek is the core author. First complete RSI framework in visual coding, one of the most subjective spaces in ML.
The core model is a Diffusion Language Model, not autoregressive. It sketches the whole scene first and refines it globally in passes, instead of writing token by token and getting stuck with early mistakes.
Real-time-level latency. SOTA visual code quality. 1/20th the DiT cost.
This isn't incremental. It's the video substrate getting quietly rewritten in front of everyone
@GeekParkHQ @SigmaZAI_office Hey GeekPark! Great to see you here! Thx for the comment! 👍 I’ve been following your early vc investments in china’s tech, great track record, well done!