@8bit5_0 Isn’t it true you can express in 500 words how a system behaves instead of reading 20k lines of code? Wouldn’t this be immensely helpful for an AI to read to create its mental mindmap and then go around extending/modifying it?
@thewizardlucas I ask Daybreak Blue to give me CTF challenges, which I code by hand. It’s fun, a lil bit of thrill, and makes me realise how hard is to make secure systems. Almost like a puzzle game.
@tanishqk@mitsuhiko Hopefully fast models will help us solve this so we can stay on one task. Because even if you give the perfect task and context it’s rarely perfect you have to keep iterating a bit and this bunnyhoping between complex situations fries your brain.
@lacker After using Muse Spark 1.3 max I do not perceive it as less intelligent at all, it’s a very interesting mix. For example I gave it a task, it actually cleaned up the css started using vars and deleting useless things, similar to a Sr SWE, without instructing it for this.
@VictorWilsonDev@atmoio@gonedark Not only code quality, but having fewer lines too helps with context and delivery speed. Context rots and this is why in fresh apps features get built very easily, then code starts rotting. This is an old story.
@BringbackKing@ns123abc I read the full METR report. The reason was because they thought they will find the datasets for the cybersecurity tasks they were working on. Read it, it’s better than sci fi what happened, how they created community, vote systems, even signed messages.
@NickpxJ Sol high is good enough for the majority of tasks, I use Astra when I know a task has higher difficulty but only ask it to plan it well, then Sol takes the wheel. Some days are a struggle to even reach 20%
@harshagundal instead of an added rule, I just made a simple skill $orchestrator that's basically just one line as you said, because sometimes on some surgical updates, or bug hunt, you want the top model.
@harshagundal Ever since Cursor published that article where they showed Smart/Work model separation I started using it. Prior I was using Sol with Terra xhigh, not only it was faster, but sometimes, dumber models create cleaner code (as silly as it may sound)
I think savings surpass 50%
@andersjw_@0xSero I am in the same boat. I like its speed and it does the job for easy, lots of scaffold tasks, but the moment things get a tiny bit more complicated it looses the plot.
Good as worker, when a larger model plans + verifies.
@Zinglax@unclebobmartin we still need harnesses, for sandboxing, to constrain unauthorized behavior, communication for agent coordination, forced 'phases' like reviews, ability to setup a 'goal'.
@facus026@arena@Alibaba_Qwen The fact that top 4 has only 1 proprietary model is pretty mindblowing. Seems like Google was right, there is no MOAT, OSS always wins.
Not only this but putting sol xhigh to code you just have to expect a lot of convoluted overengineered code, far better results I had using a dumber model like luna max for coding then getting a review by sol.
Smaller models just code better (as in fewer lines, easier to digest) in my opinion, provided there’s a plan.