Visual content is getting its source code back.
AI can now write a visual program, render it, inspect the result, and revise the source. The output is no longer just pixels. It is an editable, executable artifact.
We call this Visual Code. https://t.co/Rer6VbZldD
AI makes a lot of things faster. It doesn't make the thinking optional.
Hallucinations are hard to catch not because they're absurd, but because they're smooth. The logic closes. The tone is certain. Every word sits exactly where it should.
Some things just don't go faster. What I do now: write the ugly version myself first, then bring AI in.
Flip that order and you stop being the author. You're just the proofreader.
A lot of AI video looks rough. Doesn't matter. It moves fast enough to enter while the trend is still hot, sometimes fast enough to start the next one.
Same with AI writing. Mediocre, zero soul, and for most low-threshold demand, good enough is simply good enough.
The pattern: AI-generated stuff doesn't have to win the traditional quality game. Quality was never its card. Matching each person's need is.
Toutiao caught on-demand distribution. This wave of AI products might be catching on-demand generation.
Completion rate might be the most overrated metric in content. All it really measures is how many seconds people tolerated you.
The moment content can be questioned, expanded, rewritten, the data changes entirely. Where viewers stop to ask. Which direction they branch toward. That's intent, not tolerance.
"How long they watched" and "what they wanted to see" are different animals.
One is a feedback loop. The other is a stopwatch.
Was thinking about why "interactive video" never took off. It's been around for a decade: pre-placed hotspots, pre-recorded branches, click to jump. Bandersnatch was probably the peak.
And then it hit me: that whole genre is still pre-recorded content. Just many copies of it. The branches were hardcoded. Your choice only picked which copy to play.
The actual dividing line, I think, is whether the interface and the branches get generated at the moment you act. No hotspot placed by a human, no branch recorded in advance.
One is multiple choice. The other is a conversation. Kind of weird that they share a name.
Something I keep noticing about mass media: information basically flows one way, past the user. Newspapers, radio, TV, short video. The formats changed, the direction never did.
You can pause, skip, like. But you can't ask it "why," and you can't make it explain again differently.
AI might be the first time mass-distributed content can respond to the person watching it. The only real precondition is that generation and rendering costs drop low enough.
Once the cost line is crossed, interaction stops being a feature. It becomes the medium.
Realized there's a real difference between a generated image and a generated visual program. One is an output. The other is an asset.
An asset is something you can edit, version, drop into a pipeline, re-render under different conditions. An image just sits there.
Production never cared about the moment of generation anyway. All the value is in what happens after.
Been thinking about why visual code generation sits directly on the test-time compute curve.
It turns a visual problem into a verifiable coding problem. Code, render, inspect, revise. Every loop gives you a precise signal about what to change.
The model isn't generating more images. It's debugging a visual program inside a renderable environment. That's where compute actually converts into quality.
Iterating with diffusion models is basically re-rolling dice. Generate 20 images, pick one, try again. The feedback is global and fuzzy.
Code-native iteration is a different animal. Spacing wrong? Edit the CSS. Curve off? Fix the path. Animation feels sluggish? Adjust the timing. Every round improves the artifact itself, not just another sample.
One is sampling. The other is debugging. They converge very differently.
Pixels vs. code
For years we judged visual AI by one thing: how good the pixels look.
But I'm increasingly convinced that's the wrong yardstick. A designer doesn't want a mockup, they want layers and components. An animator doesn't want a video, they want keyframes and timing curves.
Users never wanted the final output. They wanted something they can keep changing.
Following this thread: once immediacy is solved, games and video both become personalized generation, and disposable 3A titles stop being absurd.
The real watershed is the weakness you named at the end: when models can natively see and play the worlds they build, rendering, auditing and iterating collapse into one loop, and the jank fixes itself. At that point content is no longer manufactured, it's performed live, like improv theater: when the audience leaves, the world dissolves.
Just came across this a16z piece. It lands almost exactly where we'd already landed on our own.
"The most interesting visual AI tools today have stopped trying to generate the final output. Instead, they're generating the source code behind it."
That's Tap8.
https://t.co/jAi57cvqmr
The future of AI content is hybrid. Pixels win on texture and atmosphere. Code wins on structure, text, data, timing, and anything you'll need to edit tomorrow.Explainers. Product demos. Data stories. Motion graphics. That's all structure.
That's all code.
Spent Saturday at a friend's place. A few of us ended up watching founder interviews in his living room, and it turned into a long conversation about AI and models.
The line that stuck with me: the stuff worth learning from has a short shelf life. Current opinions, current playbooks, current cases. Classics are for first principles, not tactics.
Same era, same variables, so you can hold it up against your own business and see which part to borrow and which part to fix.
Starting a company is practice, not an exam.
The #1 thing I want from any model right now: speed
Intelligence is already good enough. But waiting 1-5 min per task is the worst possible window: too short for deep work, too long to just stare at the screen. So I end up scrolling X
Agents are making all of us more ADHD
We ran a head-to-head benchmark of 7 AI video systems: Tap8, Fable 5, Seedance 2.0, Google Veo 3.1, and others.
The clips in this video are pulled directly from that benchmark so you can see the difference yourself.
She had a question. The creator wanted to answer it, and couldn't, not for her and not for the thousands of people asking their own version of it.
That's true of every creator, every video, every question left hanging.
DLM is what makes it possible for the video itself to answer, live. Vision-guided RSI is what makes it keep getting better every time it does.
Today, we're releasing the official Tap8 launch video.
Today we shipped the world's first research preview of a DLM-based interactive video model.
Benchmarks: SOTA quality and latency, unlocking a brand-new AI product form factor.
On practical video generation, measured by hundreds of human evals + VBench, SigmaZ outperforms Fable 5 and Seedance on quality, stability, cost, and speed. New SOTA.
Join Waitlist:https://t.co/jAi57cvqmr
Every AI video you've watched does one thing: plays.
Tap8 doesn't just play. Click something in frame, ask it a question, ask for something different. It generates a real answer, live, right then. Not a pre-recorded option. Not a script.
Why that's actually hard — and what it looks like — is worth watching end to end.
@william_ya51538 @lawhcd walk through it below.