"How to evaluate?"
Many people have asked me, so I finally wrote a blog post about it. I believe that the stronger the evaluation is, the easier the road to a good product is. Our benchmark system shows us the quality vs speed trade-offs for any model!
https://t.co/4qFGDgjNte
Pretty cool! Claude 5 Opus outperforms Claude 5 Fable on our bug finding benchmark (but I'd say that mainly the cost is nice)!
We're still running our automatic hillclimbing on our other workflows (since otherwise it's not fair!), so curious about those results...
I never have to wait to test a time-based app again...
Traveling time is supported in our automatic end to end tester!
Can't wait to see what crazy time-based apps it will need to test from users...
@khorrami_amin I checked if people have sold subscriptions for stroopwafels in Amsterdam, luckily seems like not too many have died: https://t.co/yAdeElqa2r
When a new business joins Biscuit, we start by understanding what they do, then suggest what's genuinely worth building for them. Take @ClickHouseDB: below, you can see how a trial desk tool gets built for them in a matter of minutes. Fully end to end.
This might be interesting: https://t.co/bpLUVOe8ev
Can External Validation Tools Improve
Annotation Quality for LLM-as-a-Judge?
Also check its references, lots of work in this area (and Iβm sure newer work too). I think there are a few domains where this makes a lot of sense, but keep in mind that we want to keep the judge cheap and fast.
This was fun!
I spoke about "scenario driven development" at the AI Native meetup, and the discussions were fun!
You can read more about our evaluation setup here: https://t.co/4qFGDgjNte
Gave a lecture about pre-training, continued pre-training (or mid-training) and when I would not try to train any parameters at all!
This was at the Computer Vision Seminars and guest lecture of the Foundation Models course at the University of Twente that @NicStrisc organizes.
It's always great to speak with him!
@simon_x_wang@Biscuit_so Hmm, good question!
I think there are a few existing benchmarks that do this. TriviaQA is the traditional "academic" one, but e.g. GAIA is closer to what you mention.
GAIA is also included in BenchLM (https://t.co/5AqsmFRGjJ), which I like.
Gemini 3 Flash vs Gemini 3.5 Flash (internal @Biscuit_so benchmarks)
Spoiler: 3.5 Flash is noticeably faster...
...but 230% more expensive on our view creation benchmark
Same prompt, same design system and tools as input for the models:
Opus 4.7 vs Opus 4.8 (internal @Biscuit_so benchmarks)
Spoiler: Clear gains on editor & exploration tasks, but mixed elsewhere.
Same prompts β here's how they differ visually π
On our billing safety benchmark (finding billing bugs, compliance issues & risky edge cases):
β 76% (4.7) vs 74% (4.8) (basically equal)
Worth upgrading for editor & exploration work? Early signs are promising.
What are you seeing with Opus 4.8 so far?