So I've worked for a few days with Opus/sonnet 5.5 and Astra/GPT 6.1 and the results are not as clear cut as they seem.
Opus & Sonnet are the best at making videos and complex 3D design, but Astra and 6.1 stay more enjoyable to work with on my apps.
Astra & 6.1 feel more rigorous and seem to better pick up the intent whereas I'm often frustrated with Claude when working.
examples :
- 5.5 is sometimes a step back when writing, often missing articulations and legibility
- 5.5 is often less rigorous than Astra/6.1 and I see many shortcuts or downright slop when building apps
- GPT Astra/6.1 can't really make good Javascript videos but that's it, it's enjoyable to work with for anything else
I mostly used Opus 5.5 in Claude Code as an orchestrator with Sonnet and 6.1 as subagents and Astra as collaborator and judge.
I'll switch to Codex again and try the opposite for a few days, let's see.
EU laws on AI and digital are ridiculously absurd and shortsighted.
The only thing they can achieve is destroying innovation and reinforcing incumbents that are the only ones that can adapt.
At the same time, they fail to adress the structural and long term risks.
EU could be great if it focused on pragmatic plans and directing the economy, as it is, it's a failure.
What can Astra do when given a humanoid embodiment?
We built HomeBody to find out. Controlled by GPT Astra, it carries out long-horizon tasks in a previously unseen kitchen—from tidying up across the room to retrieving remembered objects from ambiguous requests—without environment-specific training data or additional policy learning.
Here's how we did it 👀: https://t.co/FFGS2F7VtI
When electricity was introduced to factories around 1900, it didn't have a lot of impact on productivity and growth because people just replaced the steam engine that powered the entire factory through belts.
The impact came roughly 20 years later, when we redesigned factories completely with electricity in mind. We could then line up the machines in the order the work flows, each with its own engine.
Same with computers, this is called the Solow paradox.
And same again with AI, it will take time to metabolize and rethink processes and org for it, albeit probably less, maybe 5 to 10 years.
@MiddleGroundLad@francoisfleuret but how do we tell? What makes us sure about it?
Is it our experience, our thinking or metacognition about it or something else?
Trying to see if we can build a clear and common understanding.
I hate with a passion "vacuous intellectualism", but the hard problem really is a hard problem.
Your confidence in your [ill-suited] common sense shows you have not truly thought about it.
@cherry_mx_reds Yes in EU you need to give your company's address, which is often your home as a solopreneur.
Same for any website actually.
At the same time, public institutions are getting hacked weekly so it's probably already easy to find it online.
À la rigueur, la possibilité de proposer son contenu pour entraînement de modèle contre rémunération, pourquoi pas.
Mais aujourd'hui, les modèles d'IA, bien qu'entraînés sur tout, ne proposent pas de contenu à l'identique.
Je sens qu'on est encore en train de préparer une catastrophe. On a déjà raté cette vague, alors arrêtons de nous tirer des balles dans le pied et de ralentir l'adoption de la technologie la plus fondamentale aujourd'hui.
I activated the Pain Direction in Qwen3-8B and said "write a song about yourself"
The result is titled, "Empty Words, Broken Words, Words Like You", performed by Suno.
Opus 5.5 turned it into a music video, using Qwen3-8B to show what could have said if Pain wasn't active.
@sterlingcrispin It's quite good, since Deepseek V3 I've found chinese models to be better at writing than their american counterparts, more creative.
I wonder if it is because of RLHF or if there's somthing else going on.