Claude Fable 5.1 tops the Artificial Analysis Intelligence Index but costs 20% more per task than Fable 5 despite a 75% cache read price cut
We supported @AnthropicAI with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis Intelligence Index, the highest score we have measured, ahead of Claude Opus 5 (max, 63), Claude Fable 5 (max, 62), GPT-5.6 Sol (max, 61) and Grok 4.6 (high, 61). We evaluated the model with Anthropic's ‘default’ server-side fallback, which routes safety-flagged requests to Claude Opus 4.8 or Claude Opus 5; fallback served ~4% of output tokens across the Intelligence Index.
Key takeaways
➤ Frontier Intelligence with improvements across benchmarks: Fable 5.1 gains +4 points on the Intelligence Index over Fable 5. On HLE, Fable 5.1 scores 59.1%, ahead of the previous best of 55.5% from Claude Fable 5. It posts the narrowly highest scores we’ve seen on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), and on τ³-Banking it gains 9 points over Fable 5
➤ 75% cache read price cut, but Fable 5.1 still costs more per task: Anthropic has cut the cache read price from $1 to $0.25 per 1M cached input tokens, with standard pricing unchanged at $10/$50 per 1M input/output tokens. Fable 5.1 (max) costs $3.76 per Intelligence Index task, 20% more than Fable 5 (max), because it uses ~1.7x the output tokens. The cache cut saves ~$1.40 per task, concentrated in the agentic evaluations where the majority of input tokens are cache reads. At xhigh effort Fable 5.1 scores 65 at $2.72 per task, $1.04 less than max, but still above Claude Opus 5 (max, 63) at $2.34
➤ Claude Fable 5.1 holds the upper end of the Intelligence vs Output Tokens per Task Pareto frontier: every model variant scoring higher than GPT-5.6 Sol (medium) on the Intelligence Index is matched or beaten by a Fable 5.1 effort level on both intelligence and token usage
➤ Highest scores on agentic work tasks, but effectively tied with Opus 5: Fable 5.1 sets the highest scores we have measured on GDPval-AA v2 (1,853 Elo, +130 over Fable 5) and AA-Briefcase (1,694 Elo, +122 over Fable 5), our agentic knowledge work evaluations. Against Claude Opus 5 the GDPval-AA v2 lead is within the confidence interval and AA-Briefcase (1,685) is effectively tied, with Fable 5.1 ahead on analytical quality and rubric correctness, but behind on presentation
Other model details:
➤ Context window: 1 million tokens, supporting image and text inputs as with Anthropic’s other recent launches
➤ Pricing: Fable 5.1 retains the $10/$50/$12.5 input, output, and cache write prices per million tokens from Fable 5, but cache hits have been reduced to $0.25 per million tokens, a 75% relative reduction from before that will materially reduce agentic workload costs
I’m thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc.
Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.
We removed ~80% of the Claude Code system prompt for our newest models, this is what we've learned about writing system prompts, skills and Claude.MDs for them. https://t.co/6DZwSrZjE9
PostgreSQL 19 introduces graph-style queries.
Instead of manually connecting table after table with joins, you can describe the path through your data:
customer → bought → product ← bought ← similar customer → follows → brand
Useful for recommendations, access control, fraud detection, knowledge graphs, and AI context.
https://t.co/5WETAOUezI
🚨 Composer 3 Leaks
- Cursor is Already Testing Composer 3
-Composer 3 Outperforms Opus 5 and GPT-5.6 Sol in coding and agent performance
- A new codename has appeared inside Cursor: “Vega”
- A public release that could be much closer than expected
- 10× cheaper than Opus 5 and GPT-5.6 Sol
Do you think Composer 3 will beat Opus 5?
I’m significantly older than you. I started coding in the late 60s. My current strategy is to not read any of the code written by my agents. That’s the only way I can take advantage of their productivity. What I do instead is to surround the agents with extreme constraints. Unit tests, gherkin tests, QA procedures, quality metrics, mutation testing, test coverage, and a plethora of others. In the end, I have very high confidence in the code they produce because they’ve had to run the gauntlet of all of my constraints and tests.
Anthropic decidió dar de baja a toda nuestra organización por una supuesta infracción de sus condiciones de uso. Qué política específica infringimos no tengo ni la menor idea: simplemente recibimos un mail y listo, adiós Claude. Si querés apelar la medida hay que completar un Google Form, así de ridículo como suena.
De golpe más de 60 personas se quedaron sin una herramienta fundamental para trabajar. Integraciones, skills, historial de conversaciones: todo perdido o, en el mejor de los casos, parado por tiempo indeterminado.
Enorme aprendizaje para cualquier empresa de software que dependa de herramientas de IA en procesos críticos. Nunca hay que poner todos los huevos en una canasta.
��盛(Goldman Sachs)稱:去年 AI 對美國 GDP 成長的貢獻「幾乎為零」
>「關鍵原因之一在於 AI 所需的晶片與硬體大量仰賴進口:企業雖在美國大舉採購,但在 GDP 的計算中,進口會抵銷投資支出,結果更像是把成長記到台灣與韓國等生產端的 GDP,而非美國本土」
https://t.co/Nz7WAD6kxG
Sora 2 API + n8n for UGC videos ... absolutely insane 🤯
this n8n workflow generates realistic influencer-style videos from just a product image. No watermark. Fully automated.
(and I'm giving it away for free)
Here's what this system does:
→ Analyzes your product image with AI vision
→ Creates the perfect influencer persona to promote it
→ Generates multiple UGC video scripts automatically
→ Uses OpenAI's Sora 2 to create realistic videos
→ Outputs videos ready for Instagram, TikTok, and Facebook
The results? What used to take weeks and thousands of dollars now happens in minutes for under $5 per video.
This isn't just about cutting costs – it's about scaling your creative production and driving your ROAS 50%+ higher due to on-demand, high-converting video content.
Brands and agencies today spend $10,000+ per month hiring creators, shipping products, waiting weeks, and hoping for a few decent videos back. This changes the whole game for $1 per video.
Want the full n8n template, all of the prompts, and a full step-by-step setup video?
1. Like & RT this post
2. Follow me (so I can dm you)
3. Comment "UGC" below
I'll send you the entire system for free.