@thsottiaux@TJeparskis Yes it's cheaper per task, but i think we give astra way more tasks than sol, since it's so good! So usage just burns away so quickly
@ZechCodes@ArtificialAnlys You're right, and that means that we can't just look at the intelligence index when choosing a model. It would be like choosing a coworker based only on IQ: it misses many important things.
But apart from that: these tests are still not even measuring capability that accurately.
@victornunez Yeah i agree, but it's harder and slower to show off than one-shotting a game. It potentially requires days / weeks of collecting data, while one-shotting a game requires minutes / 1 hour.
@lyraxana Wow, so much better than 5.6 sol could have done. And this model's strength is probably not 0-shots, but continual iterations for long time, so this is even more impressive. Also Astra's attention to detail is insane, its "mind" feels so large
@luisvelasco Model architecture, data, how you made your models so efficient. What progress have you made towards generalization and continual learning.
@Bitcoin_Teddy Wtf, is he happy running a black box binary generated by ai by the end of this year? Not only is today's AI still too dumb for that, but even if it were capable, how can you trust such an output, if you can't even inspect the code?
@ananayarora Actually a very myopic take. As others have pointed out, if it costs 20k $ NOW to find that bug, it's gonna cost WAY less by the end of the year, with probably open source models being able to do it. And in a few years the cost will be close to 0. So yes, it changes everything.