@JaroslavBeck 3.What is your intuition on how this would behave with smaller and larger models (e.g. 4B or 200B+ MoE)? Do bigger models overthink more, so the savings scale up?
4.LiveCodeBench actually improved. why does coding benefit from shorter reasoning?