Good example of the difference between Sol and Fable:
I was optimizing a system by evaluating with a benchmark.
Sol made good progress towards increasing the score.
Fable realized sol had been optimizing for the bench and created an opaque eval so the improvements generalize.
@ex0t1clol@OpenAIDevs Flex inference takes ~5-10 mins per turn or so. So if modeled after that it would be dramatically slower but still useful for specific things.
@thsottiaux Would it be possible to add webhook support to codex similar to scheduled actions right now? Can basically do it already by polling but webhooks would be better.
@pa1ar Does anyone know how well dynamic reasoning effort works on luna max? Is "max" just the highest it will go, or a recommendation, or some combination?
@paulg@Austen Do you think weaker models will eventually get good enough for most use cases and become broadly commoditized? Right now the major labs’ biggest moat is just having the best model
In computer architecture work is done optimizing for specific benchmarks because they’re what reviewers use. The problem is 99% of real users do nothing like the benchmark workloads.
imagine a benchmark of two editors
one opens a file 10x faster! wow it must be better. but oh wait they just disabled syntax highlighting
real products have to do things that make it worse on benchmarks but better in practice
@pmarca “Make your answers as long and detailed as you possibly can” is surprising. I find it’s easy to ask for more but you often want a quick overview first.