I was inspired to cook up this epic rap battle vid between two of my favorite folks, @poteto (Grok Bot) and @thsottiaux (ChatGPT Dots).
Play it with sound on! Who do you think is winning right now? 😅
@xavier_mitjana Yo recomiendo Grok Bot, más aún si tienes cuenta en Cursor. Por US60 tienes un desarrollador de software con IDE más un Muktiagentes como GrokBot. Es maravilloso lo que he podido hacer con esto. Rinde muchísimo.
Eleven v4 is INSANE! 🤯
Here's my 44-sec spec ad for it.
I wrote a script of everything an AI voice "can't do." Then made Eleven v4 read it out loud.
Sound on 🔊
used sol 6.1 for a whole day and i'm very.. confused
will try my best to break down the key things to know
1. sol 6.1 reminds me of opus 5
this is a model that makes me constantly have to ask "wtf are you talking about"
it's very capable. most of the time once i understood what it meant, it did mean the right thing. but it's a pain to get through its choice of words
i hope there's a sol 6.5 coming that ends up being what opus 5.5 was to opus 5 and get this fixed
2. it's bigger than previous sol
you can easily observe the slowness that comes with this model. this hints at it being a bigger model than 5.6 sol. but my confusion is why it's priced at terra and sonnet range
then i realized, the pricing is set to attract enterprise customers who pay at API rate
for consumers who buy the $200 subscription, the lower price doesn't mean more usage because they're cutting the subscription's quota by half
so i guess the intention is to give enterprise customers an opus competitor at lower price, while holding back the benefit from consumers
3. all gpt models still suffer from poor judgment
sol 6.1 is likely taught by astra. if you read some of its code, it smell like astra code. it also inherited astra's strength which is that it can get stuff done with less turns and tokens (even opus 5.5 hasn't got to this level of efficiency yet)
but... the biggest problem to me is still "judgment". this is most obvious when i start to use either astra or sol 6.1 as my firstmate - everything simply starts to fall apart because of poor decision making
for example it would see me approving a plan and decide to file it in the backlog rather than the more obvious next step which is to start implementation (which opus 5.5 and fable would always "just know" it should do)
all signs point to a lack of human judgment in post training - all gpt models are likely heavily optimized by RLVR
overall - i'd categorize sol 6.1 as a solid implementer and reviewer, but i would not use it interactively
openai has work to do - right now they simply don't have an opus 5.5 equivalent anywhere in its line up, which is problematic
but seeing how anthropic could come back from opus 5, there's still hope :)