@Im_IrushiK Codex and Sol. Honestly, Sol is amazing when you design with 5.6 Pro and then implement with 5.6 Sol on xhigh and don't use sub-agents or ultra. I had it start one other Codex thread (not a sub agent) and work together on building. Truly Amazing what I got done in a day
@jun_song I can hand Sol twice as much as 5.5 so I do. And guess what, it takes twice as much credit to do twice as much work. Start paying attention to your work and go back to 5.5 if you want to see the difference
@jun_song Sol doesn't burn limits any faster than 5.5. It's been tested and confirmed a dozen times. People just building way more and using ultra which is just running a dozen sol instances in parallel for a 20% better result
@KaviAntonio@shiri_shh That's actually false information. Just a rumor that has been floating around but that is at least explicitely denied by the company. Composer is built on Kimi (which they got a license for). Inkling is a full new pretrain for their own dataset
@AstraiaAI That isn't necessarily the best sources but if it's true that explains why they are so darn expensive... Amazing how poorly they perform for their parameter count honestly
@TechByTaraa VS Code is a pretty great editor... And also I hate it and want some real competition. I literally built my own as a multi-repo got manager first and then an editor second and it's pretty nice
@rajgupta7813589 Wow... Antigravity is actively trying to separate itself from VS code and VS-code likeness and Codex is just adding more and more VS-Code like features over time
@Web3Aible@iruletheworldmo Idk that either of you have tested any of these models in the things this bench is designed for π
Benchmarks are good and useful but also test specific things and can have their answers leaked. So exactly which bench matters a ton (e.g. don't trust SWEbench pro anymore at all)
@geluhorotan@TheSpacerr Gemini is used accidentally by almost everyone... And now they don't have the compute to train models with. Also they tried to train with TPUs which are great for inference but terrible for training. I think that probably put them behind on speed a bit too
@TheSpacerr Google used ai for optimizing web search and then hopped into LLM side of things for consumer use late. Turns out user data is more valuable than the entire Internet (also bad teamwork and some gambles that didn't pay off)
@shiri_shh Oh also Don forgot Mercury 2. It's not a top intelligence model but it is the fastest model with useful intelligence. Cerebrus GLM4.7 and Groq Kimi K2 are the only two who actually compete with Mercury 2 at instant results currently
@noor36758 Haven't used Fable but I suspect 5.6 Pro fills the gap (I use it for the architecture side of things and it's a huge improvement over an already amazing 5.6 Sol)