The Marionette Test: DeepSeek V4.1-Flash Max | (5m 00s / $0.04)
130x cheaper than Fable. Five live strings, a real cascade with the hands taking turns, and the only entrant that admits someone is up there: two strings run off the top of the frame instead of a bar hovering in the void. But the hands barely move, and it set the loop to half the pattern's true length, so a ball teleports across the chest every 1.5 seconds.
The Marionette Test: Astra vs Fable vs. Muse Spark vs. Gemini
A marionette only works if the bar above moves the limbs below. Juggling only works if the balls obey gravity and the hands take turns. The model has to hold both in its head and write out the coordinates without ever seeing what it made. Although not a frontier model, we included Gemini 3.8 Flash so it feels included and won't form a complex it takes into adulthood 🫂
Claude Fable 5.1 Max | (24m 27s / $5.80)
Best of the four. Correct gravity, best lighting/shadows, the only one that built a proper two bar rig. But it drew four strings, its legs never move, and the top bar isn't connected to anything and nothing is holding it up.
GPT-6 Astra Max | (17m 42s / $2.03)
Nice suit, plenty of strings. Its near knee bends backwards. Gravity wise it looks like he's in outer space, so the balls hang like it's juggling on the moon.
Muse Spark 1.3 Max | (4m 42s / $0.18)
Most complete rigging, seven strings running to head, hip, both knees and both hands. Open curtains and a city skyline (best background other than Fable IMO). Middling physics, rough suit, but at least it has real arms unlike Gemini.
Gemini 3.8 Flash Max/Extended Thinking | (1m 18s / In-App)
Despite being a Flash model, Gemini has the best juggling animation outside of Fable. The balls also have faint shadows on the ground, which GPT-6 Astra failed to do. The only one that worked out that a strict side view flattens a cascade into a single plane, and it corrected for it. Good gravity. But the arms are bare sticks inside a sleeveless suit and the background is empty.
Nobody made the bar drive the puppet.
People don't like to accept the fact that intelligent folks are also subject to GroupThink. X is not a welcoming place for anything other than AI accelerationism.
There's also the giant shadow of financial interest fueling a lot of the fervor against pausing / slowing down.
A lot of otherwise smart people on Twitter seem 100% convinced AI risks are all fake and stupid and part of some marketing ploy. What is surprising is that some of these people seemingly also believe that AI’s positive uses are on an incredible trajectory of increasing capability with no end in sight. VC Twitter is particularly infected by this pattern. It’s not really coherent.
Most positive use cases for AI have a corresponding “dark version”. If you are super human at coding, you are also super human at hacking. If you are superhuman at structural engineering you are likely superhuman at finding structural flaws to knock buildings down. If you are superhuman at designing drugs, you are superhuman at designing novel undetectable poisons. If you can cure viruses, you can create them. Some of these “dark versions” are not so bad, and some are actually pretty scary. Either way these are real societal and technical problems that need to be solved to get the good stuff and avoid the bad stuff. We’re experiencing the first of these with coding and computer security which is the most advanced, but that won’t be the last. I think we’ll be able to solve these problems, but they aren’t solved yet and if you believe in continued AI progress they are surely coming.
But “bad people using AI” is not the only problem. Uncontrolled AI autonomously doing bad things, despite sounding kind of nutty, is also something we should be concerned about. AI “killing us all” is not the most likely outcome, but the chance of a major civilization-wide catastrophe doesn’t have to be very high for it to be a concern. Again, this is only a problem if capabilities advance to a point that AI can do really crazy things on their own, which hasn’t happened yet, but I think the Hugging Face incident is a good example of the outline of how things can go wrong when capabilities outpace alignment. We should be glad that the only available bad thing right now is hacking, which isn’t all that bad. It’s clear that as you get to superhuman capabilities you need a level of alignment and control that is correspondingly superhuman. Humans have plenty of misalignment problems themselves (serial killers, mass shooters, tyrants, etc), but it’s a manageable problem because most humans can’t do that much damage and we’ve developed systems to prevent dangerous humans from getting too much power.
Talking about these issues is just common sense. It’s not a sign of some kind of neuroticism or pessimism. These are just hard problems that it’s very important to solve for AI to have a positive impact. I’m pretty confident we will solve them. But we haven’t solved them yet, and to my mind we are clearly on a trajectory of rapidly increasing capabilities which means this is important. Mocking people who are worried about this or talking about it, without anything substantive to say about how we can be sure these problems won’t arise, is not really a very helpful contribution.
Today I made the costly mistake of using Claude extra usage instead of just signing up for an entirely new plan. And ended up spending $90 on roughly 4 or 5 hours of work
Its interesting watching this community of Highly Educated Free Thinkers rush to a conspiracy that the "CIA is running a psyop to turn people against AI" because they can't fathom reasonable people being worried about this. 🍿
Jacob Coxon's post on X is over 120M views, he's everywhere. He's in TIME, Variety, the WSJ, on NBC and Fox. Christiano to the OpenAl board. Daniel Kokotajlo on Rogan. OpenAI just called for Congress to create mandatory national safety regulations. The wind has suddenly changed.
@sophiamyang@Muse I purchased something last night with Muse completely through guest checkout, no login.
I guess it depends on where you're buying from
Here's a few reasons why people have so much trouble predicting the future:
1) Fail to realize that even the best superforecasters' predictions drop off dramatically past a three year time horizon and that most folks are worse than dart throwing monkeys at future predictions.
2) Fail to realize you literally can not see black swan inventions coming around the corner. If you predict the future of Germany in 1439, then in 1440 your predictions are completely wrong because of the Printing Press.
If you're an 18th century farmer you can't see a web developer job because it exists on the back of countless developments and inventions you can't predict.
3) They change one variable and hold all other variables the same. i.e. AI advances and nothing else does, no parallel discoveries or innovation, no solutions, no mitigations, no societal or cultural changes.
4) Predict unlimited resources and zero friction in the real world (dust, disconnects, diffusion, etc) to slow/divert/change/impact the development. All changes experience equal and opposite reactions.
5) They mistake their ability/expertise in a domain for a parallel/orthogonal ability to predict the future of that domain and its impact on the world. Two different skills and they do not usually overlap (though very rarely they do.)
6) What I call "classic sci-fi or Jules Verne syndrome", which is similar to one variable changes. It's like in Jules Verne when one guy gets the submarine and nobody else does. But life is more like cell phones, lots of people getting them over time in a diffusion curve.
7) They mistake exponential curves as infinite always and never see an S curve coming.
8) The see infinite resources (compute/memory/learning upper limits/improvement) and no limitations.
On one hand, you get a really cool AI assistant that helps organize your email and shop.
On the other hand, a 10% risk of AI yeeting humanity into oblivion during one of the most fragile moments in the American empire. Very tough choice.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
The Marionette Test: Astra vs Fable vs. Muse Spark vs. Gemini
A marionette only works if the bar above moves the limbs below. Juggling only works if the balls obey gravity and the hands take turns. The model has to hold both in its head and write out the coordinates without ever seeing what it made. Although not a frontier model, we included Gemini 3.8 Flash so it feels included and won't form a complex it takes into adulthood 🫂
Claude Fable 5.1 Max | (24m 27s / $5.80)
Best of the four. Correct gravity, best lighting/shadows, the only one that built a proper two bar rig. But it drew four strings, its legs never move, and the top bar isn't connected to anything and nothing is holding it up.
GPT-6 Astra Max | (17m 42s / $2.03)
Nice suit, plenty of strings. Its near knee bends backwards. Gravity wise it looks like he's in outer space, so the balls hang like it's juggling on the moon.
Muse Spark 1.3 Max | (4m 42s / $0.18)
Most complete rigging, seven strings running to head, hip, both knees and both hands. Open curtains and a city skyline (best background other than Fable IMO). Middling physics, rough suit, but at least it has real arms unlike Gemini.
Gemini 3.8 Flash Max/Extended Thinking | (1m 18s / In-App)
Despite being a Flash model, Gemini has the best juggling animation outside of Fable. The balls also have faint shadows on the ground, which GPT-6 Astra failed to do. The only one that worked out that a strict side view flattens a cascade into a single plane, and it corrected for it. Good gravity. But the arms are bare sticks inside a sleeveless suit and the background is empty.
Nobody made the bar drive the puppet.
@s_batzoglou I do. Anthropic is serving the model they have the compute the reliably serve while being aligned. Not the most powerful model they technically can make.
So I guess it depends on what you mean by "Ahead". Because if you can't bring your best stuff to market then...
@ryanlpeterman@trq212 Interesting things in this interview:
- Model system prompts are getting simpler, but harnesses are only getting more complex. H2 of 2026 is definitely the era of the harness
- The concept of giving the model permission to burn tokens and ambitiously complex its work
@Adam__Allcock@reach_vb@steipete When you have unlimited budget you have the gift to experiment, when you have the gift to experiment you can learn lessons / encounter pitfalls faster than others. I'd love to hear about rabbitholes he went down that ended up being a waste of time / or look promising
@Prathkum We should stop thinking of these models as the best things the labs can produce, but rather the best model they can reliably serve and be aligned at the same time.