Had to add gemini-3-flash-preview to the results.
It dominates.
Clearly the top model on this benchmark.
Hopefully we can get a v2 of this bench out sometime soon.
Long context eval.
Huge improvement since last year. The frontier models went from poor to great.
An exciting standout is kimi-2.5. It made impressive progress without (presumably) a new architecture, putting up gemini-2.5-pro numbers which we were all impressed by last year.
Long context eval.
Huge improvement since last year. The frontier models went from poor to great.
An exciting standout is kimi-2.5. It made impressive progress without (presumably) a new architecture, putting up gemini-2.5-pro numbers which we were all impressed by last year.
@k0tovsk1y Been working on a better one for the past few months, hope to get it out soon. But at the same time these models are just now good IMO, and you'll start seeing that in the real world in terms of agentic workflows starting to work frfr.
claude-opus-4-5 fixed claude's long context performance, it is now good when previously it was a laggard. claude-sonnet-4-5 had a regression compared to sonnet 4โฆ Same tier as grok-4.
Kimi-k2.5 now the Chinese/Open-source leader!
Minimax???
gpt-5.2 improves on almost perfection in gpt-5 to now very close to perfect. gpt-5.2-pro did surprisingly poorly.
@teortaxesTex I might, but my bench is saturated by gpt-5. It's meaningfully better than gemini 2.5 and the bench did not reflect that. I will be back with a better eval.
Fiction.LiveBench for Long Context Deep Comprehension adds: deepseek-v3.2-exp [reasoning: high], deepseek-v3.2-exp, nemotron-nano-9b-v2:free, qwen-max, qwen3-next-80b-a3b-instruct.
Some updates to Spiral Bench:
- A more detailed rubric for protective vs delusion-reinforcing behaviours
- Responses evaluated by a judge ensemble: sonnet-4.5, gpt-5 & kimi-k2
- New models evaluated: qwen3-235b, glm-4.6, grok-4-fast, mistral-medium-3.1
Thoughts: Interesting that we see an improvement for deepseek's reasoning mode but no improvement for the non-reasoning. It has high scores on the easier questions but very low scores on the hard ones.
grok-4-fast is fairly close to sonama-sky-alpha while still being free.