Everyone is hyping up fable 5.1 and astra... but are they actually that much better?
YES they ARE.
saw a sample of fable 5.1's output compared against gpt-5.6 sol and opus 5
yeah prompts were different and idk how many shots/retries fable needed, but the detail is actually insane
timeline rumors say gemini 3.8 flash, astra, and fable 5.1 are all dropping this week
we’ll see if big tech corps can reclaim the hype or if open-weights already cooked them beyond repair lmao
This voxel comparison got way more interesting than I expected
I gave Claude Opus 5, GPT-5.6 Sol, and Gemini 3.7 Flash the same basic challenge: turn a reference image into a large, interactive voxel scene of a Japanese temple complex
And somehow, all three created completely different worlds
Claude built the tallest and most architectural version, with a huge central pagoda and a dense, polished temple district
Sol created the largest environment: 67,097 voxels, multiple temples, bridges, ponds, a waterfall, cherry blossoms, and atmospheric night lighting
Gemini took a cleaner and more minimal approach, with brighter colors, a simpler garden layout, and a strong central pagoda
Same reference. Same idea. Three completely different interpretations
I didn’t write any code. I just shared the reference, explained what I wanted, and gave a few follow-up adjustments
I put all three results into one 50-second vertical video so you can compare them side by side
Which one would you choose?👀
Prompt and reference image are in the comments
Just tested Gemini 3.8 flash in antigravity on the same one-shot 3D
I'm not impressed at all
glm-5.3 flash is STILL the clear winner here, and i haven’t even tested fable 5.1 or opus yet lol
solid 6/10. yeah it’s cheap, fast af, and textures improved over 3.7, but food assets still look like low-poly generic blocks. UI works but the overall visuals are completely mid
definitely gotta run astra vs fable 5.1 on this benchmark next fr
This is actually INSANE. Can’t believe how GLM-5.3 (ex-0xAlpha) won and AGAIN beat corp models like grok, gemini
look at this API pricing per 1M tokens:
- input: $0.075
- output: $0.25
- cache: $0.015
costs literal pennies and passed the one-shot vibe check way better. solid 7/10
UI is cleaner, and while 3D food assets still look like pure shit, the lighting, logic, and zero texture clipping are actually working
but nobody said this is THE benchmark, it’s literally just a vibe check
on top of that my whole point is there’s no perfect model, every model’s good at some stuff and trash at other stuff (and read the post, i literally said glm’s a W here in that specific case lol)
don’t take it too seriously, it’s just one test outta thousands on that platform. i just like checking what ai can visually pull off rn, not really interested in like how fast it rewrites an excel file or whatever
@Ananth7e feels like it was straight up benchmaxed fr
and i’m not even talking coding here, just asked some linux terminal questions and it took way more steps to fix the issue than it shoulda. didn’t have this problem with 3.7
@bridgemindai is anyone even surprised tho? kinda obvious google and tough coding tests just don’t mix
p.s. still think it’s the best model for the average user tho
OMG. this is actually insane. fable 5.1 just completely cooked every other model
have you guys seen the outputs yet?
that same rocket launch prompt everyone’s been benchmarking: bridgeminde ran it through his setup and it is straight up PURE CINEMA fr
vfx look crazy, but the sound design is the real diff = easily the best audio generated by an ai coding-model
idk what astra can even pull off to compete with this lmao, bar is in the stratospher
WTF... 0xAlpha is actually on another level
just threw the same rocket launch one-shot prompt at it and I’m speechless
easy 9/10. and this is ALL ONE PROMPT on a FREE model (for now)
honestly can’t even find any real flaws here, maybe some slight overexposure on the launchpad model but that’s literally it
everything else is butter smooth, features work, audio goes crazy
the AI race is moving way too fast bro
Everyone is hyping up fable 5.1 and astra... but are they actually that much better?
YES they ARE.
saw a sample of fable 5.1's output compared against gpt-5.6 sol and opus 5
yeah prompts were different and idk how many shots/retries fable needed, but the detail is actually insane
timeline rumors say gemini 3.8 flash, astra, and fable 5.1 are all dropping this week
we’ll see if big tech corps can reclaim the hype or if open-weights already cooked them beyond repair lmao
This voxel comparison got way more interesting than I expected
I gave Claude Opus 5, GPT-5.6 Sol, and Gemini 3.7 Flash the same basic challenge: turn a reference image into a large, interactive voxel scene of a Japanese temple complex
And somehow, all three created completely different worlds
Claude built the tallest and most architectural version, with a huge central pagoda and a dense, polished temple district
Sol created the largest environment: 67,097 voxels, multiple temples, bridges, ponds, a waterfall, cherry blossoms, and atmospheric night lighting
Gemini took a cleaner and more minimal approach, with brighter colors, a simpler garden layout, and a strong central pagoda
Same reference. Same idea. Three completely different interpretations
I didn’t write any code. I just shared the reference, explained what I wanted, and gave a few follow-up adjustments
I put all three results into one 50-second vertical video so you can compare them side by side
Which one would you choose?👀
Prompt and reference image are in the comments
@Ananth7e ngl i test out a bunch of chinese models regularly, they’re pretty solid too, like glm 5.3 flash for example. but noticed if i gotta tackle something actually tough, fable or opus is my go-to number 1 pick every time
This comparison got way more interesting than I expected
I gave grok 4.6 build, gpt-5.6 sol, gemini 3.7 flash, and glm-5.3 flash the exact same prompt + image ref: build an interactive arcade water sim with sharks in three.js
all 4 cooked up completely different outputs:
❌ gemini 3.7 (biggest L): output runs fine and tap works, but it completely ignored the reference image aesthetic and built it in a totally different style lol
🌊 grok 4.6: reflections look like a high-budget render and beach matches the ref. sharks look a bit scuffed tho, and ripple physics are way too aggressive
💀 glm-5.3 flash: water looks simple but realistic, sharks are decent, but tap is buggy, and it casually melted 1.5M tokens lmao
🔥 gpt-5.6 sol: by far the closest to the image ref and prompt. tap is mid, but the rest is straight fire (not a pure one-shot run ngl)
zero perfect models exist...
each has its own pros and cons, and anyone claiming ai can reliably one-shot complex stuff rn is on pure copium