Think Big App, Small Model: GPT Luna vs. Terra vs. Sol
Defaulting to biggest model can 5x to 10x AI Cost.
๐ Think Big App, Small Model
๐ Model escalation is the exception
๐ Continuously flatten the complexity curve of your system
#ProductionCoding
https://t.co/jSLa5LdTtb
@AjaySinghadiya8 I've used v4-pro for production coding (#rustlang) for months.
It's good, but Luna is close, if not better.
Those v4-... models reason a lot, and reasoning costs $$
Luna's speed is a game changer
I'll try the new v4-pro, but not for "open weight", since it's no longer cheaper
Significant price increase expected
DeepSeek's cheap days may be over
DeepSeek v4-flash was already pricier than Luna on real tasks
It'll be interesting to see v4-pro's price
I am aaig DeepSeek fan and user, but OpenAI Luna my first choice for now (mostly #rustlang coding)
@Da7_Tech but 4x to 8x slower than Luna.
Luna got webgpu/sharder of my benchmark first go.
v4-flash, did not.
but still good to have those open models putting pressure.
@westxsh@theo Yes, way overdue.
Next one will be about 3 Levels channel pattern
Level 1 - app wrapper api (on top of crossfire)
Level 2 - type alias on top of level 1 (e.g. TuiTx, TuiRx)
Level 3 - type wrapper on top level 1 (e.g JobQueue, ...)
Hoepfully this weekend ...
@mattpocockuk To be honest, I never got this grill-me thingy.
I do
chat md -> goal md -> plan md -> managed loop
Then put/update spec mds with goal/plan when needed
Goal, plan, loop are optin
@Da7_Tech Luna is cheaper or same price and so much faster.
Been a big fan/user of dflash for months. But Luna new pricing is game changer.
the Luna speed for similar dflash price is crazy
Qwen 3.8-max is good,
but so SLOW, and NOT cheap
(from qwencloud .com)
For this demo, WebGPU, 2 prompts (see in the comments)
output: 63k tokens (including 56.5k tokens)
price: $0.40
time: 20 min !!!
Luna-max was $0.06, 1.5 min, for a similar result
๐ prompt in comment ๐
Playing with WebGPU / Shaders for my next AI benchmarks
Kind of cool, but not really what I asked for.
Tuning...
Had to use Sol to finish Kimi K3 work.
K3 did good though... but so slow ...
@Da7_Tech Agreed. With the right flow, the price structure is a game changer.
I've been doing some pretty big stuff with Rust, and it's surprisingly good.
And it's so nice to have a 10-run loop for $0.30.