First benchmark with Union Alpha — the Sunset Ocean prompt.
This is a pretty demanding visual test: procedural ocean movement, reflections, lighting, atmosphere, clouds, and keeping the entire scene visually coherent while animated.
Union Alpha did a solid job overall.
The water is easily the strongest part — lots of surface detail, good wave variation, and the sunlight reflection reacts nicely across the moving ocean. The horizon stays clean, and the scene maintains good visual consistency throughout the animation.
The weaker area is the sky. The clouds feel a little flat/heavy compared to the water, while the sunset could use richer color separation, softer atmospheric scattering, and more depth around the horizon. The sun/reflection also feels slightly too bright and uniform in some areas.
Overall, a good result for the first benchmark.
Water/physics: Strong
Lighting/reflections: Good
Atmosphere: Decent
Clouds/sky: Needs improvement
Visual coherence: Strong
What do you think about this result?
More Union Alpha benchmarks coming.
Grok roadmap is looking pretty serious.
Grok 4.7 → roughly Opus 5.0 level, better in some areas and worse in others. Multimodal still needs work.
Grok 4.8 → noticeable improvement over 4.7.
Grok 4.9 → expected to reach Astra/Fable class.
Grok 5 → potentially better than anything available today.
The jump from 4.7 → 5 could be huge.
@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.
Grok 4.8 will be a noticeable improvement.
Grok 4.9 is probably Astra/Fable class.
Grok 5 maybe better than anything. We shall see.
Sam Altman just said the main thing he was excited about launching THIS WEEK has been pushed to NEXT WEEK — and that it’ll be “worth the wait.”
Interesting.
We literally just got GPT-6 Astra two weeks ago, and OpenAI already appears to have another major launch lined up.
No confirmation yet on what it is, so I’m not going to pretend we know.
But if this was the release Sam was most excited about this week…
Next week could be interesting.
Some early feedback on Union Alpha after testing it:
The model itself is genuinely good. I’ve been pushing it through some pretty demanding coding and graphics tasks, and I’m impressed with what it can do so far.
My biggest complaint right now is speed.
It’s VERY slow.
Some generations take long enough that it noticeably affects the testing experience, especially on larger coding tasks.
That could simply be because everyone is testing Union Alpha right now and the infrastructure is under heavy load. We’ve already seen how much usage this model is getting, so I wouldn’t be surprised if that’s the main reason.
But purely as feedback: latency is currently the biggest thing holding back the experience for me.
Quality looks promising. I just want that quality a lot faster.
Early Union Alpha feedback:
The model is genuinely good. I’ve been pushing it through some demanding coding tasks and I’m impressed so far.
My only complaint is speed — it’s VERY slow right now.
Could simply be the massive usage from everyone testing it, but latency is definitely the biggest issue for me.
Quality is there. Just need it faster.
I’ve been testing Union Alpha pretty heavily today.
Instead of throwing small coding prompts at it, I’m running it through a set of demanding real-world graphics and agentic coding benchmarks designed to expose where a model actually starts breaking down.
So far, I’m testing three major scenarios.
Test 1 — Photoreal Real-Time Ocean Rendering
A complex Three.js/WebGL graphics task where the model has to build a photorealistic sunset ocean scene from scratch.
This isn’t just “make some animated water.”
I’m testing whether it can correctly reason about and implement things like physically plausible wave simulation, custom GLSL shaders, water BRDFs, Fresnel reflections, procedural detail, atmospheric scattering, realistic sun reflections, clouds, post-processing, camera motion, anti-aliasing and performance optimization.
The final result also has to remain performant at 1080p/60 FPS.
Basically: can the model behave more like a real graphics engineer instead of just generating something visually impressive for one screenshot?
Test 2 — Constraint & Iteration Test
I’m also running another variation of the graphics benchmark to see how consistent Union Alpha is across repeated complex builds.
This matters more than people think.
A strong coding model shouldn’t occasionally produce an incredible result and then completely fall apart when given the same level of difficulty again.
I’m looking at instruction adherence, architecture decisions, shader quality, visual quality, runtime stability and whether it actually implements the difficult requirements instead of quietly simplifying them.
Test 3 — Full Rocket Launch Simulation
This one is significantly harder.
Union Alpha has to build a cinematic orbital rocket launch sequence entirely inside Three.js.
That means procedurally creating the vehicle, launch pad and environment while also handling a shader-driven rocket plume, smoke and steam particles, changing exhaust behavior with altitude, lighting, atmospheric effects, max-Q condensation, stage separation, second-stage ignition, broadcast-style camera work, telemetry, procedural audio and a deterministic launch timeline.
Again, everything has to run inside a single self-contained browser experience without relying on downloaded assets.
This benchmark is interesting because it tests several capabilities simultaneously:
• Long-horizon instruction following
• Large codebase generation
• Three.js/WebGL knowledge
• GLSL shader programming
• Physics reasoning
• Visual taste
• Performance optimization
• State/timeline management
• Debugging ability
• Whether the final project actually works
I’m deliberately making these tests difficult.
Modern models are getting so good at small coding demos that “build me a cool website” doesn’t tell me much anymore.
I want to know what happens when the specification becomes long, interconnected and unforgiving — where one bad architectural decision can affect ten other systems.
Union Alpha has been extremely interesting to test so far.
I’m still running the benchmarks, so I’m not giving it a final judgment yet.
Once all three runs are finished, I’ll share the actual outputs, what it implemented correctly, what it skipped, where it struggled, and how it compares with the other frontier coding models I’ve tested.
This should be a much more useful test than another leaderboard number.
I’ve been looking into Union Alpha, and this might be one of the most interesting stealth model launches recently.
Here’s everything I’ve found so far:
Union Alpha appeared out of nowhere as an anonymous model focused heavily on coding, research, and agentic workflows.
• 262K context window
• Multimodal input
• Tool calling
• Agentic coding support
• Currently FREE
• Available through OpenRouter/OpenCode
• Provider still unknown
What caught my attention first was the coding performance.
Union Alpha reportedly scored ~74% on DeepSWE, putting it around some seriously strong frontier coding models on that benchmark.
But I’m even more interested in the scale behind this launch.
The model reportedly crossed 100 BILLION tokens while serving around 500M tokens per minute.
Then capacity was pushed even further — toward 1 BILLION tokens per minute.
Scaling a stealth model to that level of inference before we’ve even figured out who’s behind it is wild.
I’ve also seen a lot of speculation that Union Alpha could be related to GLM or another unreleased frontier model.
But I wouldn’t treat any of that as confirmed yet.
Right now, the provider is still anonymous. We don’t officially know the lab, architecture, parameter count, or actual model identity.
So I’m separating the speculation from what we can actually verify.
What I know so far:
Free.
262K context.
Multimodal.
Tool calling.
Strong early coding numbers.
And an insane amount of inference demand.
I’ve started testing Union Alpha myself.
Benchmarks are interesting, but I care much more about how these models actually perform when I give them real coding tasks.
I’ll share my results soon.
Scaling a stealth model from 500M to 1B tokens/min before figuring out compute capacity is wild.
100B+ tokens served already.
Whatever model this is, people are absolutely hammering it.
This.
OpenAI and Anthropic seriously need to push these models harder into medicine and other fields that directly improve people's lives.
Solving Navier–Stokes and other Millennium Prize problems is incredible, and I’m all for it.
But if we have models this capable, the real goal should be turning that intelligence into better healthcare, faster discoveries, cheaper treatments, and technology normal people actually benefit from.
OpenAI and Anthropic seriously need to push AI advancement in medicine and other fields now
what is the point of having these extremely capable models if normal people don't benefit from them
Solving Navier Stokes and other millenium prize problems is incredible and im all for it, but it doesn't change most peoples lives at all
🚨 CONFIRMED IN MY TESTING: GPT-Astra Medium can somehow burn MORE usage than xHigh.
I tested both reasoning levels myself.
Medium drained my allowance insanely fast.
xHigh lasted noticeably longer.
So I dug into it — and other users are seeing the same thing.
Here’s where it gets weird:
→ Astra Medium is actually cheaper per individual response
→ Independent benchmarks also show Medium costing less per task on average
→ But on long agentic/coding runs, Medium can take WAY more turns
One reported comparison:
Medium: 122 responses
xHigh: 65 responses
Medium was cheaper each turn — but kept making more turns, reprocessing context and continuing the task.
Result?
Medium burned roughly 2× the total quota.
Another user reported exhausting a fresh 5-hour window in ~17 minutes on Medium vs ~35 minutes on xHigh.
So no — xHigh isn't magically cheaper per token.
But higher reasoning may sometimes plan the task better, avoid retries/tool calls, and finish with LESS total usage.
Which means the reasoning-effort curve for Astra might not be as simple as:
Low < Medium < High < xHigh.
For long agentic work, it can apparently flip.
OpenAI still recommends Low/Medium if you want to stretch allowance, so this either depends heavily on workload or something weird is happening with Astra's subscription metering.
Anyone able to reproduce this with the SAME repo + SAME prompt + fresh usage window?
Looks like the banked reset issue from this morning is getting fixed.
Some resets used in ChatGPT Work and Codex weren’t fully applying.
OpenAI is now giving everyone affected another banked reset + sending an apology email.
Good to see them make it right quickly.
There was a bit of a kerfuffle this morning with some banked resets not fully applying when used in ChatGPT Work and Codex. Everyone who used one in the affected time window is getting another one and an email to apologize.
While the timeline is glued to Images 2.5 and a Millennium Prize headline, Greg is pointing at something quieter and maybe more important: ChatGPT just posted its 4th consecutive monthly MAU record in August — 1.06B.
That number matters more than another feature drop because it’s hard to fake at scale. Models can look flashy on a launch thread. Four straight monthly active records is closer to “people actually building habits around this” than “people screenshotting demos.”
What’s interesting about the timing:
- Same day OpenAI ships Images 2.5 (faster gen, better fidelity, multi-edit consistency)
- Same day the lab is celebrating frontier math-agent work
- And the president is amplifying a boring-looking growth chart
IMO the tell isn’t the absolute MAU figure — it’s the streak. Fourth consecutive record month means the product is still compounding after the hype cycles everyone assumed would flatten it. That’s the opposite of benchmaxxing. That’s distribution + retention doing work models alone can’t.
Clearer-eyed read: the labs racing on capability demos are competing for mindshare, but the durable lead still looks like whoever owns the daily habit loop at billion-user scale.
🚨 OpenAI just dropped ChatGPT Images 2.5
Sam joked that it probably can’t solve super difficult math problems — but for image generation, this is a serious upgrade.
What’s new:
• Up to 50% lower generation latency
• Better reference + identity preservation
• More precise image editing
• Stronger multi-turn consistency
• Better lighting, textures, layouts, and style adherence
• Sketch directly inside ChatGPT
• Image comments for targeted edits
• Templates for posters, merch, product shots, and more
OpenAI is also launching two API models:
• GPT-Image-2.5 Flare → faster, built for scale
• GPT-Image-2.5 Sunburst → higher precision for serious creative work
OpenAI says users are already creating 3B+ images every week across ChatGPT Images + its APIs.
IMO the biggest upgrade isn’t raw image quality.
It’s finally being able to say:
change this ONE thing
…and have the model actually preserve everything else.
That’s the difference between an image generator and a real creative tool.
Images 2.5 feels much closer to an AI-native Photoshop workflow.
ChatGPT Images 2.5—faster, sharper, smarter, with better tools for creating whatever you can dream of.
- Faster image generation to keep your ideas flowing
- Improved fidelity for more natural, recognizable images
- Consistent details across multiple edits
- Comment-based edits to change only what you want
🚨 OpenAI just dropped ChatGPT Images 2.5
Sam joked that it probably can’t solve super difficult math problems.
But for actually creating and editing images, this looks like a major upgrade.
OpenAI says more than 3 BILLION images are already being created every week across ChatGPT Images + its image APIs.
And Images 2.5 is focused heavily on one thing:
making image generation actually usable as a creative workflow.
What changed:
• Up to 50% lower generation latency vs Images 2.0
• Better preservation of people, pets, products, and reference subjects
• More natural lighting and richer textures
• More precise image editing
• Better consistency across multiple rounds of edits
• Improved understanding of complex visual instructions
• Stronger style adherence
• Better handling of complicated layouts
• Improved transparent-background generation
The editing improvements might be the biggest part.
One of the biggest problems with current image models is this:
You ask:
“change this ONE thing”
…and the model decides to regenerate half the image.
Images 2.5 is specifically designed to edit the requested element while preserving:
• Subject identity
• Composition
• Background
• Previous edits
• Brand treatment
• Overall visual structure
And that consistency is supposed to hold across multiple editing turns.
OpenAI also added new creative tools directly inside ChatGPT:
Sketch
You can draw a rough idea directly inside ChatGPT and use it as a reference.
For example:
• Sketch a room layout
• Draw the placement of objects
• Rough out a clothing silhouette
• Mark a composition
• Add basic visual structure
Then describe what you want and Images 2.5 turns that rough sketch into a finished image.
There are also:
• Image templates for posters, merch, product photography, and more
• Comments you can place directly on an image for targeted edits
• Shareable prompts so other people can remix workflows with their own images
The API release is interesting too.
OpenAI is launching two GPT-Image-2.5 models:
GPT-Image-2.5 Flare
The faster/default option.
OpenAI positions it for:
• Social content
• Creator tools
• Product experiences
• Visual search
• Rapid prototyping
• High-volume image generation
It targets higher quality than GPT-Image-2 while cutting latency by up to ~50%.
Then there’s:
GPT-Image-2.5 Sunburst
The higher-precision option for more demanding creative workflows.
Best suited for things like:
• Production-ready ads
• Campaign creative
• Product photography
• Brand assets
• Repeated precision edits
Basically:
Flare = speed + scale
Sunburst = control + quality
And this isn’t a tiny limited rollout either.
Images 2.5 is rolling out across:
• ChatGPT Free
• ChatGPT Go
• ChatGPT Plus
• ChatGPT Pro
• Business / Enterprise
• ChatGPT Work
• Codex
• Web
• Desktop
• Mobile
Flare + Sunburst are also available through the API.
OpenAI is continuing to use C2PA metadata and invisible watermarking for generated images as well.
IMO the most important improvement here isn’t raw image quality.
Image models are already insanely good at generating a beautiful first image.
The harder problem is:
• Generate something
• Inspect it
• Point at one specific part
• Change only that part
• Preserve everything else
• Repeat this 10 times
• Without slowly destroying the image
That’s what turns an image generator into an actual creative tool.
Images 2.5 feels like another step away from:
“AI image generator”
and toward:
“AI-native Photoshop / creative workspace.”
If the multi-turn editing consistency is actually as good as OpenAI claims, that’s a much bigger upgrade than another benchmark jump.
Going to test this hard.
OpenAI just claimed a solution to the Navier–Stokes Millennium Prize Problem.
And the most interesting part isn’t another “AI does math” headline.
It’s the system they used.
According to OpenAI, Sam Altman, and Greg Brockman:
a next-gen model, described as significantly more capable than GPT-6 Astra
~10,000 coordinating AI agents
~88 hours of work
Lean formalization built into the process
all while the underlying model is still training
The result: an analytical proof arguing that Navier–Stokes dynamics can produce a finite-time singularity, with the fluid vortex spiraling inward and becoming increasingly elongated.
If this holds up, that alone is historic.
But there’s another important layer here.
OpenAI went out of its way to address provenance and attribution.
They credited Levent Alpöge and Tristan Buckmaster, said the agents and researchers did not see their work before its public release, said no specific user data was accessed to solve the problem, and acknowledged that they still can’t completely rule out indirect influence from de-identified product-usage data.
That matters.
Because after the attribution drama around AI-generated mathematical discoveries, the real frontier isn’t just:
“Can the model solve it?”
It’s:
“Can you prove where the result came from?”
Benchmaxxing a Millennium Prize tweet is easy.
Coordinating ~10,000 agents into a Lean-checked mathematical argument in under four days is a much scarcer signal.
And doing it while publishing the provenance caveats alongside the capability claim is arguably just as important.
Sam called this one of the most amazing moments in OpenAI’s history and said he didn’t expect a result of this magnitude this soon.
The broader implication isn’t “AGI solved fluids.”
It’s that scientific discovery may increasingly look like:
frontier model
→ agent swarm
→ automated research
→ formal verification
→ human scrutiny
The labs that win science probably won’t just be the ones with the smartest models.
They’ll be the ones that can prove both the theorem and the chain of custody.
Capability can win the headline.
Formal verification + attribution hygiene will decide who scientists actually trust.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
DeepSeek is moving FAST.
1M context + vision, the API beta is already live, and now V4.1 Flash is showing up in Vercel AI Gateway.
Feels like the full V4.1 release is basically around the corner.
September is getting ridiculous 😭
🚨 DeepSeek V4.1 Flash just appeared in Vercel AI Gateway
• 1M context
• text + image input
API beta went live this a few hours ago too
V4.1 looks basically ready to drop in the next few days once this beta wraps up