why do anthropic models almost always return empty JSONs/responses when benchmarking for minebench ππ
i'm like 95% sure it's because max reasoning effort exhausts the output budget, but it happens so often idk if it's an adapter issue atp
The model appeared defeated, spending the next several hours doing essentially nothing but farming potatoes. Stream viewers noticed and complained that Astra needed to "pick up the pace." Astra also seemed to become paranoid about creepers: "GREEN tall thing ahead was SUGARCANE, NOT creeper!" At times it was hard on itself: "you can screw up and drop things"; "do NOT waste another night chasing dark pink pixels" (its own disparaging wording for pigs); "our last tool stupidly ended Slabsselected/rightUP and 15sec thinking killed us".
@PHPLego@developedbyed nit, but it's not really the "MineBench voting system"; LMSYS' Chatbot Arena popularized the blind, head-to-head voting system for LLM evals :)
@R2Cdev_@minebench_ai Top is launch day, bottom is today
there being no obvious difference was my point; if Astra had actually been "nerfed," you'd need a consistent drop across repeated runs, not one worse sample
we also don't know if OP kept the same prompt/harness
FWIW, here are some MineBench builds Astra generated on release day and today with the same prompt
One-off before/afters don't tell you much about stochastic models; if there even was a change, it could just as likely be from changes in the harness
(top: release, bottom: today)
Astra (launch day) vs Astra (Today).
Same prompt. Today's result looks off.
The photo realism is gone. It clearly looks nerfed.
GPT-5.6 Sol is unusable for me right now too.
anyone know if Codex banked resets change your weekly reset date?
like if i have a weekly reset in 3-days and wanna use a reset to get some work done before then, or would it also extend my weekly reset date to 7-days out?
i've always felt happy with my typing speeds but with how fast agents have become, i feel like typing not only wastes time but is also less effective at conveying what you want
@Dimillian@cherry_mx_reds can i ask if there's a reason you've stuck to default context settings?
i've kept a ~350k compaction window, but with astra not billing extra over 250k tokens, and improving so much on MRCR, idk if there's a reason to not increase it closer to 1mil π€
@minebench_ai started as a side project... 400k+ blind votes, 300 GitHub stars, and nearly 300k page views later, itβs pretty surreal to see whoβs finding it
*cough* @gdb iβm graduating May 2027 and looking for new-grad engineering roles π *cough*
This is personally one of our favorite benchmarks!
Inspired by MazeBench, we had GPT-6 Astra build us a maze on MineBench, then tried to solve it ourselves :)
We couldn't make it through, but maybe you can: https://t.co/pupv6CVaui