Deepdive: Kimi K3 max vs. GPT 5.6 Sol max on software engineering/DeepSWE tasks.
Both of these models provide nice specializations and routing/cascading between them is the right way to go!
full deep-dive 👇(1/n)🧵
@_Suresh2 true but its still worth routing see the numbers below - even a non-optimial router will save money and get better perf https://t.co/jI8JmqlV3A
Kimi K3 and GPT Sol are pretty different in their behaviour and you can use that to route and/or cascade between them
depending on how good your router is you can get upto a nice boost over either model while still keeping costs lower than Sol alone👇
I had an LLM classify each software eng task into categories like data modeling, runtime internals, testing and tooling etc.
then I looked at the performance for each category for Kimi K3 vs GPT 5.6 Sol
if you can classify your task you can build a fairly optimal router using this youself!
How to route, you ask!?
Classify tasks by what the code is: Sol leads 5 of 8 domains, Kimi 3.
Sol should be selected for: serialization (92-79), concurrency (72-55), program analysis (64-56).
Kimi K3 is better for: ops tooling (79-73), runtime internals (77-75).
Conformance: a 61-59 coin-flip.
I had an LLM classify the task given the prompt from the benchmark!
I've been doing a lot of summarizing, labelling, classification for thousands of long context model trajectories and i must say the speed and intelligence per dollar value for @MiniMax_AI M3 is pretty unmatched.
this is the distribution of programming languages in SWE Bench + (multilingual) + SWE Atlas + DeepSWE + Terminal bench
1357 tasks total
> lots of python, go
> Rust and C are pretty low
> this proportion does not line up well with how popular the underlying language is
@realbrucemartin Yeh also i think heuristic based routing(route by programming lang, task type classification etc.) might just be good enough
Ppl get bogged down in training the optimal router which is hard
@xcate329@togethercompute what I meant was you could route between the two or use a cascade strategy where you always go to kimi first and then fallback to Sol on failure
detailed perf boosts shown here https://t.co/jI8JmqlV3A
Kimi K3 and GPT Sol are pretty different in their behaviour and you can use that to route and/or cascade between them
depending on how good your router is you can get upto a nice boost over either model while still keeping costs lower than Sol alone👇
The graph suggests open model usage > closed.
But openrouter usages def. leans more into open and free models.
I think closed model usage is still much larger than open but open models are definitely starting to make a dent.
Long way to go still.
@ajs6888 The more different the model performance the higher the chances that you can mix and match between the models to outperform any single model strategy.
I suspect a bunch of labs would be trying to poach from deepseek:
"There's only one thing we can't compromise on: maintaining team stability. This was a very significant risk we faced. Of course, this risk has been largely mitigated with this round of financing."