Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
We supported @AnthropicAI to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far. Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59), Kimi K3 (57), and Claude Opus 4.8 (max, 56)
Key takeaways:
➤ New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max). On AA-Briefcase, our proprietary agentic knowledge work benchmark, it scores 1720 Elo, +146 ahead of Fable 5. These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup
➤ Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA
➤ Frontier intelligence with reduced cost: Claude Opus 5 (max) costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53. However, at high and xhigh reasoning efforts Opus 5 can outperform both Opus 4.8 and Claude Sonnet 5 at a lower cost per task
➤ Frontier agentic terminal use: 89% on Terminal-Bench v2.1 at max effort, roughly in line with the leader, GPT-5.6 Sol (xhigh)
➤ Outperformance on scientific reasoning: Along with leading agentic performance, Claude Opus 5 scores 53% on Humanity’s Last Exam in line with Fable 5; on CritPt, a frontier physics evaluation developed by Argonne and UIUC researchers, it also matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra
➤ Factual knowledge still lags Fable 5: As expected from the models’ size classes, Opus 5 still has lower factual knowledge on AA-Omniscience than Fable 5. It improves +7 points on AA-Omniscience Accuracy over Opus 4.8, but answers more often when uncertain - its hallucination rate rises +14 points to 50%
➤ Improving efficiency, but only on the Intelligence vs. Cost per Task Pareto frontier at high Intelligence levels: Opus 5 outperforms Fable 5 at lower cost, but at lower effort levels it sits just behind the GPT-5.6 family on the Intelligence vs. Cost per Task frontier
Other model details:
➤ Context window: 1 million tokens (equivalent to Opus 4.8)
➤ Pricing: As with recent Opus launches, tokens cost $5/$25 per million tokens of input/output; cache pricing remains at a 25% premium for cache writes ($6.25 per million tokens) with 5-minute time to live, and 90% discount for cache hits ($0.50 per million tokens)
➤ Five effort settings (low, medium, high, xhigh, max), and support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled
🦔Microsoft canceled its internal Claude Code licenses this week after token-based billing made the cost untenable, even for a company with effectively infinite cloud resources. Uber's CTO sent an internal memo warning the company burned through its entire 2026 AI budget in just four months. American AI software prices have jumped 20% to 37%, and GitHub (owned by Microsoft) is dropping flat-rate plans for usage-based billing across its products.
My Take
The AI subsidy era is ending in real time. The same company that put $13 billion into OpenAI and built the Azure infrastructure powering most of Anthropic's compute just looked at the bill from a competitor's coding tool and decided it was not worth paying. That is not a productivity failure on Anthropic's end. Token-based pricing is forcing every enterprise customer to confront the actual cost of running these models at scale, and the number turns out to be far higher than the flat-rate experiments suggested.
This ties directly to my Gemini Flash post yesterday. Anthropic, OpenAI, and Google all raised effective prices in the last six months. Enterprises that built workflows assuming AI costs would keep falling are now watching annual budgets evaporate in months. Two outcomes look likely from here. Either enterprises scale back AI usage to fit budgets, which slows the revenue ramp the labs need to justify their valuations ahead of IPOs, or the labs cut prices and absorb the losses, which makes the unit economics worse at exactly the wrong moment. Both paths land in the same place, the numbers stop working, and somebody has to take the writedown.
Hedgie🤗
Si hoy tuviese que aprender Claude Code, arrancaría por esto:
1) Agent loop: entender cómo Claude piensa, ejecuta acciones, verifica resultados y corrige errores mientras trabaja.
2) Permissions & Auto Mode: approvals, auto mode y qué puede ejecutar Claude automáticamente.
3) Memory (CLAUDE.md): guardar reglas, comandos y contexto para no repetir lo mismo en cada sesión.
4) MCP: conectar Claude con GitHub, Slack, databases y herramientas externas.
5) Skills: crear workflows reutilizables para tareas repetitivas o específicas del proyecto.
6) Subagents: dividir tareas grandes en agentes separados para mantener el contexto limpio.
7) Hooks: automatizar validaciones, permisos y restricciones antes o después de acciones importantes.
8) Planning: cuándo usar /plan antes de tocar código o cambiar arquitectura.
9) Session management: usar /compact, /clear, --resume y --continue para manejar sesiones largas.
10) Rewind & checkpoints: volver atrás cuando Claude rompe algo o toma un mal camino.
11) Commands: aprender /permissions, /memory, /review, /agents y otros comandos clave del día a día.
12) Effort levels: cuándo usar think, ultrathink o distintos niveles de razonamiento según la tarea.
Y recién después iría a conceptos más avanzados:
context engineering, harness, context rot, multi-agent workflows y performance.
CURSO COMPLETO DE CLAUDE DE 4 HORAS
Esta es la guía más detallada de Claude que he visto en línea.
Guarda esta página antes de que se te olvide.
Construye herramientas.
Automatiza el trabajo.
Aprende cómo las personas construyen bots y sistemas.
🚨 EN VEZ DE VER NETFLIX ESTA NOCHE…
mira esto durante 1 hora.
Este curso de Claude AI te enseña a construir y automatizar casi cualquier cosa.
Los que lo vean hoy despertarán mañana con una skill que la mayoría no tendrá en 2 años.
Los que lo ignoren seguirán viendo Netflix el año que viene preguntándose por qué nada cambia.
Tu decisión.
Sígueme para más contenido así 🚀
this is actually insane
> be tech guy in australia
> adopt cancer riddled rescue dog, months to live
> not_going_to_give_you_up.mp4
> pay $3,000 to sequence her tumor DNA
> feed it to ChatGPT and AlphaFold
> zero background in biology
> identify mutated proteins, match them to drug targets
> design a custom mRNA cancer vaccine from scratch
> genomics professor is “gobsmacked” that some puppy lover did this on his own
> need ethics approval to administer it
> red tape takes longer than designing the vaccine
> 3 months, finally approved
> drive 10 hours to get rosie her first injection
> tumor halves
> coat gets glossy again
> dog is alive and happy
> professor: “if we can do this for a dog, why aren’t we rolling this out to humans?”
one man with a chatbot, and $3,000 just outperformed the entire pharmaceutical discovery pipeline.
we are going to cure so many diseases.
I dont think people realize how good things are going to get
Florida man sold his house in just 5 days after letting ChatGPT handle the entire process instead of a real estate agent
The AI handled pricing, marketing, showings, and even helped draft the contract
Para empezar en Ciencia de Datos no es necesario que te aprendas 200 algortimos de Machine Learning.
Solamente con estos 8️⃣ ya puedes resolver la mayoría de problemas básicos 👇
1️⃣ Regresión Lineal
The basic one. Sirve para entender cómo cambia una variable cuando cambian otras. Muy simple, pero útil para empezar a pensar en términos de modelos.
2️⃣ Regresión Logística
A pesar del nombre, se usa para clasificación. Muy común en problemas tipo sí/no: fraude, churn, aprobar o no un crédito.
3️⃣ Árboles de Decisión
Modelos que van tomando decisiones paso a paso, como si fueran reglas. Son fáciles de entender y de explicar a gente que no es técnica. Como un if...else..., para que lo entiendas.
4️⃣ Random Forest
Básicamente muchos árboles de decisión trabajando juntos. Suele funcionar muy bien sin tener que complicarte demasiado.
5️⃣ Gradient Boosting (XGBoost / LightGBM)
Si trabajas con datos tabulares, lo vas a ver muchísimo. Es de esos modelos que, bien afinados, suelen dar resultados muy buenos. Para producción son muy top.
6️⃣ K-Nearest Neighbors (KNN)
La idea es simple: mira qué ejemplos se parecen más a uno nuevo y decide en función de ellos. No suele ser muy práctico, pero te ayuda mucho a entender la distribución de tus datos.
7️⃣ Support Vector Machines (SVM)
Busca la mejor frontera posible para separar clases. Durante años fue uno de los modelos estrella en clasificación.
8️⃣ K-Means (Clustering)
Cuando no tienes etiquetas y quieres ver si los datos se agrupan de alguna forma natural.
Muchos de estos algoritmos se consideran modelos clásicos dentro de Machine Learning, aunque el abanico es muy amplio.
Encontrar un buen modelo no es muy dificil.
Lo difícil suele ser entender el problema, preparar bien los datos y saber cuándo un modelo realmente está funcionando.
👉 Proximamente os contaré que modelos suelen usarse para los problemas más típicos: Predicción de consumo, recomendadores, propensión de fuga, etc.
🚨 Este tipo muestra cómo crear apps móviles con Claude en solo 8 minutos.
Un tutorial directo donde pasa de cero a una app funcional usando IA, sin complicaciones y paso a paso.
Si estás aprendiendo IA, esto te interesa �
🚨 OpenAI acaba de publicar una masterclass de 32 páginas sobre cómo construir agentes de IA.
Explica cómo diseñarlos, coordinarlos y llevarlos a producción con ejemplos reales.
Si estás aprendiendo IA, esto te interesa 👇
Robots in China are doing it all now, even dancing on stage like pros.
Here Unitree robots doing Webster flips and are performing at Chinese-American singer Wang Leehom’s concert in Chengdu.
BREAKING: Sugars essential for life have been found in pristine asteroid Bennu samples collected by NASA’s OSIRIS-REx spacecraft. Combined with previous detections of amino acids and nucleobases, we see that life’s ingredients were widespread throughout the solar system: https://t.co/Tb3HpwZG9J
More on the study led by Yoshihiro Furukawa of @TohokuUniPR⤵️