so a new checkpoint of Ox Alpha dropped, hence the significantly better outputs people are noticing.
but can we get better service tho? it's unusable rn, too many errors :(
For those that still didn't get the memo for some reason, Ox Alpha is GLM 5.3 Flash
Also, a new checkpoint dropped internally today (Monday), confirmed by yours truly (the first to mention GLM 5.3 Flash)
This model appears to be a "descendant" of GLM 5V Turbo
They're currently running benchmark evals so release is very soon (if scores are good), likely the same day GLM 5.3 weights drop
@louszbd can we have vision on future glm models please π₯Ί
glm 5.2 has such a nice taste and really good model at all sorts of frontend, ui ux work but it's a blind artist π. it having vision would make it soooo much better and nicer to use!
Umm could you ask the provider to fix this maybe?
we literally can't push those token counts more than this because they aren't able to serve it reliably!
Honestly I haven't been able to get any meaningful amount of work done by Ox Alpha.
It's just not reliable, all threads come to a pause eventually (if they even respond to the first message without error)
"Error from provider (Console): Upstream request failed: Endpoint is unavailable."
I was liking this model honestly, the way it speaks, how it understands your intent etc, but can't get any real work done from it, the success rate of any decently sized task getting completed by it is like 15-20% for me.
also it's slow as fuck.
i see. how about offloading that work to cheaper sub agents? maybe luna?
i personally use luna for all cheap works like fan outs, small mechanical edits, some bounded work for non important parts of codebase, use it to do code reviews and write shit tonn of tests, computer use, monitoring stuff and all the places where I need an agent present/working but doesn't have to be the smartest model in the room.
there's free options too, like opencode, they provide lots of free models, make a skill to invoke those models so ur main agent can basically offload these works to other agents.
this 33% I used, I had sol high working for 2+ hours straight.
tho it was one thread only and I wasn't using as many sub agents as i usually would do the drain was less.
true, tho I don't usually let my usage downgrade my model pick for my work.
for my important projects i alwys use fable/sol.
and since I've made a multi provider setup, where I use both claude codex models together, so the cost of the biggest models are comparatively lesser too.
for this plus account testing i am strictly using codex only, but I was traveling today so couldn't get much done.
Current Status of the Codex Plus plan usage burn:
> Weekly Limit Used: 33%
> Tokens Burned: 72.2M
> Inference Used: $38.81
(Used Sol High and Luna Max models mostly)
At this rate Plus plan could give around $120 of weekly usage.
How is your burn so far?
I got y'all some outputs from "claude-melon-eap" and "claude-marshmallow-eap" π
They're really pushing 3D RL a lot, huh? The placement of buildings is very good
One thing I noticed on both models is they use A LOT of thinking tokens, as I reached <max-tokens> multiple times