🦔The IMF quietly war-gamed how AI could destabilize the global economy last December. The biggest threat they identified was income tax erosion. 66% of US federal revenue comes from individual income and payroll taxes. Every job AI eliminates is a paycheck that stops funding the government.
My Take
Every AI earnings call celebrates headcount reduction. Dimon cut 30-40% of jobs in some departments. Oracle cut 21,000 to fund data centers. Meta used AI to rank employees for layoffs. The entire pitch to investors is fewer workers, higher margins. But who replaces the tax revenue those workers generated? Corporate tax is a fraction of what individuals pay in income tax, payroll tax, and Social Security. The companies eliminating the jobs aren't picking up the tab.
So you end up in a cycle where companies fire workers to boost margins, the government loses revenue, and then the same government is expected to fund the safety net for the people who got fired. At some point a politician is going to look at the numbers and realize that the AI companies celebrating record efficiency are also quietly draining the tax base that keeps everything else running. CEOs don't bring this up on earnings calls, and analysts haven't started asking yet.
Hedgie🤗
⚠️AI prices are collapsing faster than any previous technology cycle:
The cost of AI intelligence has fallen as much in ~3 years as personal computer prices did in over ~15 years, highlighting an unprecedented pace of technological deflation.
According to Goldman Sachs, an index tracking the price per million tokens across leading LLMs has fallen more than 90% from its 2022 starting level in just ~30 months, with quality-adjusted pricing falling even faster.
Meanwhile, hyperscalers are projected to spend ~$757 billion on AI CapEx in 2026 alone, with cumulative spending through 2030 estimated at ~$5.5 trillion.
The challenge for the AI industry is that falling prices require explosive growth in usage just to offset declining revenue per unit of intelligence.
This is central to the debate now rattling chip stocks, if intelligence keeps getting this much cheaper this fast, the return on that $5.5 trillion could take far longer to materialize than markets are currently pricing in.
If that thesis gains traction, it could pressure valuations across the AI complex, since much of this year's stock gains have been built on the assumption that current CapEx levels will be justified by future returns.
The biggest risk to the AI trade may not be a lack of demand, but a world where AI capabilities become commoditized before they become highly profitable.
How can an LLM switch between low-, medium-, and high-effort reasoning? And how does an LLM learn to reason more or less?
I put together a “little” article explaining how these effort levels are implemented at inference time and during training.
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
🔴 ¡NUEVO KIMI K3 OPEN SOURCE!
Primer modelo open source en ponerse por encima de lo que eran los modelos frontera de hace unas semanas!
K3 supera a GPT 5.5 y a Opus 4.8 quedándose no tan lejos de la nueva generación de modelos 🔥👀
https://t.co/wK55kKQtjM
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model
Key results:
➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation.
➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality.
➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params).
➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04)
➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores.
➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities
Other model details:
Context window: 1M
Size: 2.8T total parameters
Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens.
Modality: Native multimodal input supports text and images, and the model remains text-only for output.
Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
Building a single 1 GW AI data center costs $37.2 billion up front. The chart from Epoch AI breaks where it goes.
$21 billion of that $37.2 billion is servers. The GPUs alone are more than half the cost of the entire project. The building that houses them, the facility, is $11 billion. Network infrastructure is $4.9 billion. Then land at $170 million and the utility connection at $160 million, rounding errors against the silicon.
Land plus power: $330 million combined, under 1% of the build. The thing everyone fights about when a data center gets announced, the land it sits on and the electricity it draws, is the cheapest part of the whole project. What costs money is the chips. A 1 GW AI data center is a $21 billion pile of Nvidia GPUs with an $11 billion box built around it to keep them cool and powered.
The skeptic reads this and says the chip number is a bubble artifact, GPU prices are inflated by a shortage and will collapse, so the whole cost structure is temporary. Partly fair. GPU pricing is elevated and Nvidia's margins are enormous. But the point holds even if chip prices halve, because the ratio is what matters. Servers depreciate on a 3-to-5 year cycle while the building lasts 20-plus, so the $21 billion is not a one-time cost, it is a cost that recurs every few years as the GPUs go obsolete and get swapped. The building is bought once. The silicon inside it is bought again and again. Annual operating cost is only $0.9 billion, which means the model is not about running the thing, it is about continuously replacing the most expensive thing inside it.
This is why the AI-energy debate is aimed at the wrong target. Every data center announcement turns into a fight over grid strain and electricity, and the power draw is real, the earlier Chevron and Microsoft deal was 2.67 GW for one campus. But the utility connection is $160 million on a $37 billion build. Power is a genuine constraint on where you can build and whether the grid holds, and it is close to free relative to the chips. The binding cost is access to GPUs, which is why the companies with the most compute are the ones that pre-bought silicon, not the ones that own the most land.
The whole structure is a bet on depreciation. You spend $21 billion on chips that are worth a fraction of that in four years, betting the AI they produce in those four years earns more than the next $21 billion you will have to spend to stay current.
A treadmill of silicon that has to be repurchased before it wears out.
I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally.
I thought it might be useful putting this together because many people asked me about my setup in the past, and I thought it would also motivate people to get started tinkering with local models for serious work (yes, things got incredibly capable this year with better LLMs and better harnesses).
So, here's a walkthrough of how to connect a local LLM to a local coding harness (could be Claude Code or Codex, which you may already be familiar with).
I also included some assessment notes that are useful as a checklist to select between and consider certain LLMs over others:
- Checking RAM usage at long contexts to see if the model is suitable for real work
- Measuring prefill and decoding tok/sec to see whether it's fast enough to not be annoying
- Making sure the model has sufficient tool-calling capabilities in theory
- Assessing whether the model can solve some more challenging tasks when used in a coding harness.
Of course, there are always more specialized tools that can squeeze a bit more performance out of things, but I hope this is a good starter kit that stays flexible; that is you can easily switch to newer models as they are released or even tap into cloud models in your familiar harness if the current ones are not sufficient enough for a given task.
Preocupante situación en la que parece que el gobierno federal ha frenado la salida de GPT-5.6, similar a como hizo con Fable, imponiendo en este caso una salida escalonada por motivos de seguridad.
¿Se ralentiza la salida de modelos frontera americanos?
Anthropic has released Claude Fable 5, the first publicly available Mythos-class model that ranks #1 in our agentic real-world knowledge work benchmark GDPval-AA
Claude Fable 5 shares the same underlying model as Claude Mythos 5, with added security guardrails for potentially harmful cybersecurity, biology, chemistry, and distillation-related queries. The release also introduces a fallback mechanism, allowing Claude Fable 5 to route flagged queries to a second model such as Claude Opus 4.8.
@AnthropicAI shared access with us ahead of public release to benchmark this model. Claude Fable 5 scores 1932 on GDPval-AA, our benchmark for agentic real-world work tasks, taking the #1 position and putting Anthropic models in 3 of the top 4 spots. The result was measured using adaptive reasoning at max effort, with Claude Opus 4.8 configured as the fallback model. Fable 5 falls back to Opus 4.8 on 2% of GDPval-AA tasks, with Anthropic stating that fallback occurs in fewer than 5% of sessions on average.
Full benchmarks for Claude Fable 5 are in progress - we will share the full Intelligence Index and publish scores on our website shortly
🔴 ¡CLAUDE MYTHOS 5 YA ESTÁ AQUÍ!
Llega por fin la nueva categoría frontera de modelos de Anthropic Mythos 5/ Fable 5.
En números es el modelo más potente jamás lanzado. Incluso mejorando las cifras de Mythos Preview que nos sorprendieron a todos hace unos meses.
Veamos 👇🧵
🔴 ¡NUEVO MODELO GEMMA 4 12B!
La querida saga de modelos open source de Google se actualiza hoy para añadir el modelo de 12B, que gracias a su nueva arquitectura rinde a la par que el equivalente de 26B de hace unos meses!
El modelo es multimodal en input: visión, audio y texto
Seven new models launching at Build: let’s go!
Reasoning. Code. Image. Transcribe. Voice.
Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models
Thread 🧵
#MSBuild
El paradigma de ingesta del segundo cerebro de Karpathy, donde tienes una carpeta donde echar tus datos en crudo y que luego un agente procesa y estructura para agregarla a una wiki o segundo cerebro, es un patrón escalable a otras tantas aplicaciones.
Yo en mi caso por ejemplo ya tenía creada una aplicación financiera que usaba un sistema similar: mis extractos bancarios, facturas, datos traídos de APIs, modelos de impuestos en pdf... todo en crudo en una carpeta. Con la idea de luego llamar a un agente que trabaje en dar orden y forma a esos datos (una única vez) para procesarlos adecuadamente de cara a que luego lo consuma una aplicación (en este caso en vez de Obsidian, un front-end).
Se me antoja como un nuevo tipo de aplicación con un patrón arquitectónico que funciona por poner en su diseño a un agente que cada cierto tiempo sale a pasear para dar orden al caos de la carpeta de datos. No es un script determinista que sepas que va a funcionar siempre igual, con lógicas encorsetadas a formatos concretos, sino que tiene la flexibilidad de comerse y procesar cualquier dato crudo que pongas en la carpeta.
Y donde además cualquier dato alimenta al sistema y lo mejora para hacerlo crecer.
Además, obviamente los agentes no sólo actúan como procesadores de esa información sino que luego se nutren de todo el sistema para poder hacerle consultas mucho más completas o hacer crecer tu aplicación con cada nuevo dato crudo que se agrega.
Estamos empezando a diseñar software alrededor de datos caóticos, confiando en las capacidades de una nueva capa agéntica. El usuario no se adapta al software sino que el software se adapta al caos del usuario.
So good
🔥 ¡NUEVO VÍDEO en el LAB! 🔥
Hoy analizamos a Claude Opus 4.8 y la incursión de Anthropic al mundo de los multiagentes con su nueva funcionalidad Dynamics Workflows!
Una mejora de los model...
[❌ SORRY, YOU'VE HIT YOU SESSION LIMIT]
...link a continuación 🤦♂️