GPT 5.6 Sol vs. DeepSeek Pro yields a p-value of 0.45, which is far from statistically significant.
They also cost the same per solved problem.
In today's run, there is not enough evidence to claim Sol has a meaningful performance edge, so the choice comes down to latency.
@powns_ai The benchmark uses fresh, private problems specifically to reduce contamination and gaming, so I donโt publish the underlying items. I do publish the methodology, scoring, and aggregate results so people can evaluate the benchmark itself.
https://t.co/v3fv8PD7CI
I ran it for a couple days but just like 1.1 it was scoring too low. Plus not a lot of people use it.
Hereโs the leaderboard from then. https://t.co/sTSBxEZJVa
I like to keep the leaderboard to under 15 models at a time and with all the constant new releases I drop models as needed.
The variance comes from the model providers. I just measure it. Look at the older posts, it beats the crowdโs wisdom by days to weeks.
I have nothing to benefit from with regards to what I report. The other benchmarks are either outdated within days or are biased by the labs themselves.
โRunning a startup is incompatible with being a full time studentโ
I made it through on both fronts. It was (barely) possible.
As the startup picked up customers and investors I stretched out classes over part time semesters.
Company: https://t.co/dGGPRbNz5C
I studied and graduated with a BS in Electrical and Computer Engineering.
@powns_ai Evals are done with solution references, not LLM as a judge.
Fable was consistently ranked higher until mid July when Anthropic overall dropped the ball with all their models.
Yes theyโre newly synthesized algorithm brainteasers.
My full suite across all models already takes 5 hours with a concurrency of 18 so we have to cap the time limits for practicality. This affects some models more than others.
When I ran the test yesterday Ox Alpha was actually one of the faster models. Timeouts only impacted 9 out of 40 responses. The rest were zippy.
You can see it here:
@YvesLoy@markasoftware_ Thatโs a good point. When OpenRouter shows more providers this could very well improve. Hence why I run tests multiple times a week.