Every AI video you've watched does one thing: plays.
Tap8 doesn't just play. Click something in frame, ask it a question, ask for something different. It generates a real answer, live, right then. Not a pre-recorded option. Not a script.
Why that's actually hard — and what it looks like — is worth watching end to end.
@william_ya51538 @lawhcd walk through it below.
Visual content is getting its source code back.
AI can now write a visual program, render it, inspect the result, and revise the source. The output is no longer just pixels. It is an editable, executable artifact.
We call this Visual Code. https://t.co/Rer6VbZldD
The part worth naming is why this compounds. A render can be executed and checked, so a person tunes the rubric instead of grading every sample. That's the property that made code and math move fastest, now pointed at visual work.
We run it as a recursive self-improvement loop: a coding agent writes the visuals, a vision model grades the render, a meta layer catches the systematic errors and retunes the rubric against human graders, and the traces that pass become the next round of training data.
Structure in code, texture from generative models. That's the bet.
all of the recent proc gen art, video editing and 3d game demos recently have made me update towards LLM coding models being better at a lot of creative work than diffusion models
As our CEO William Yang puts it: "AI's bottleneck is not intelligence — it's the bandwidth between AI and humans."
350+ outlets is a milestone, not a finish line. We're just getting started.
KTLA: https://t.co/8loG7ancGA
Something remarkable happened this week 🌎
Since we unveiled Tap8 Vision, the story has been picked up by 350+ media outlets across the US and beyond — from @AP to @KTLA, plus ABC, FOX, NBC and CBS affiliates nationwide 🙏
A thread 👇
So what is everyone talking about?
Tap8 Vision turns passive video into a live interface — click, question, and reshape it in real time. Powered by our Diffusion Language Model, at ~1/20 the inference cost of traditional approaches.
AP: https://t.co/iCOYR6E5Md
@lawhcd and i raised $X millions in 2 weeks to build an ai model for real-time interaction.
it natively outputs video, not just text.
because text is dead.
humans don’t experience the world one token at a time. we see, hear, move and interact in real time. ai should too.
today’s models are becoming very good at slow, deep reasoning. but most human-ai interaction doesn’t need another model that thinks for 30 seconds.
it needs a system 1:
fast, visual, always-on and instantly responsive.
we believe diffusion-LM are one path to building it.
in 10 years, people may look back at text-only ai the way we look at cave drawings today
real-time video is the new default interface to collab with ai.
Our bet: 1. codes-based video rather than pixel so it’s always accurate & naively interactive.
2. real-time rendering via ultra-fast, post-trained, Diffusion-LM (~1000+ token/sec).
3. vision-based RSI so it’s SOTA on quality (btw our dlm BEATS FABLE 5 on code-based video quality, check out our benchmarks ⬇️)
Every AI video you've watched does one thing: plays.
Tap8 doesn't just play. Click something in frame, ask it a question, ask for something different. It generates a real answer, live, right then. Not a pre-recorded option. Not a script.
Why that's actually hard — and what it looks like — is worth watching end to end.
@william_ya51538 @lawhcd walk through it below.
Worth pairing with the comprehension number: 84.5%, tied for first among all seven systems. Accurate content people don't actually understand isn't much of a win on its own, so this benchmark is really measuring two problems solved together, not one.
This is also where the architecture stops being a nice-to-have. If someone's about to buy something, follow a repair step, or study off a Tap8 video, whether the facts on screen are exactly right isn't a detail, it's the whole point.
She had a question. The creator wanted to answer it, and couldn't, not for her and not for the thousands of people asking their own version of it.
That's true of every creator, every video, every question left hanging.
DLM is what makes it possible for the video itself to answer, live. Vision-guided RSI is what makes it keep getting better every time it does.
Today, we're releasing the official Tap8 launch video.
This is our architecture showing up in the numbers.
Tap8 led Factual Coverage at 83.6% on information dense video, significantly ahead of five of six competitors.
Facts stay as live, executable code instead of being baked into pixels. That keeps equations, labels, and charts exact, legible, and interactive, while pixel diffusion handles how the scene looks.
Tokens for what is true. Pixels for what it looks like.
Full research: https://t.co/D26LIOMhiX
The technical reason for the gap: Tap8 doesn't route everything through pixel diffusion. Facts, numbers, labeled diagrams go through a live code layer that never passes through lossy video compression, exact by construction. Only the perceptual texture around it is diffusion-rendered.
That's also why it's fast enough to run in real time. The core model is a Diffusion Language Model, refining code in parallel across denoising steps, around 2,146 tokens/sec versus roughly 100 tok/s for autoregressive generation.
Benchmark results: 75.2 T-score, first of seven systems. 83.6% factual coverage, 84.5% comprehension, tied for first. To be fair, pure pixel-diffusion systems like Seedance still lead on raw aesthetic preference, but Tap8 is the only system in the set that's both high-informative and aesthetically competitive on information-dense content.
Full methodology: https://t.co/4jHUhZNG6b
We ran a head-to-head benchmark of 7 AI video systems: Tap8, Fable 5, Seedance 2.0, Google Veo 3.1, and others.
The clips in this video are pulled directly from that benchmark so you can see the difference yourself.