@ElementalReason It is a valuable contribution to the philosophy of science; however, it is framed as if it constituted a fundamental discovery in physics.
El Ministro @fedesturze revela que los gobernadores tienen la potestad de bajar el costo de medicamentos a un quinto comprando directamente en el exterior y se pregunta por qué casi ninguno se animó a hacerlo: “Yo hablo mucho con los gobernadores de que ellos podrían bajar fuertemente el costo de los medicamentos importando. A mí me llama un poco la atención que ninguno de los otros gobernadores se haya tirado así a la pileta. ¿Por qué? No sé, no quiero especular. Pero lo que sí quiero es comentarlo y que lo que quede claro es lo siguiente: un gobernador podría bajar el costo de los medicamentos en su sistema de salud provincial a un quinto comprando directamente en el exterior”.
🚨🇺🇸🇦🇷 | TRUMP SOBRE MESSI: "Hoy estamos encantados de recibir a los campeones de la Copa MLS 2025, Inter Miami... y es un gran privilegio para mí decir lo que ningún presidente estadounidense ha tenido la oportunidad de decir antes: ¡bienvenido a la Casa Blanca, Lionel Messi!".
Ahora cualquiera puede ver si tu ciudad te faja en impuestos🔥
El Ministerio de Economía lanzó este Portal. Se empieza a terminar la caja negra de los impuestos
https://t.co/ZYbvEf91VI
This is the biggest news from today’s GPT-5.2 launch.
Forget the benchmark charts OpenAI showed. Forget the 100% AIME score and the SWE-Bench Pro numbers. The real story is buried in a single data point from ARC Prize: 90.5% accuracy at $11.64 per task.
A year ago, hitting 88% on ARC-AGI-1 cost an estimated $4,500 per task. Today, 90.5% costs $11.64. That’s 390X cheaper in 12 months.
Look at that leaderboard chart. The efficiency frontier is getting redrawn every few weeks. GPT-5.2 Pro, Grok 4, Gemini 3 Deep Think, Claude Opus 4.5, all stacking on top of each other in a diagonal line from bottom-left to top-right, each one obsoleting the economics of what came before it.
Here’s what most people don’t understand about this benchmark.
François Chollet designed ARC-AGI in 2019 specifically to resist brute-force scaling. The whole thesis was that LLMs just pattern-match training data and would fail catastrophically on novel abstract reasoning. Each puzzle is unique, never seen online, requiring genuine generalization from minimal examples. Humans solve 95% of them easily. For years, the best AI systems couldn’t crack 5%.
The 2020 Kaggle competition topped out at 20%. By 2023, still only 33%. GPT-3 scored literally 0% via direct prompting. The AI research community largely accepted ARC-AGI as proof that scaling alone wouldn’t reach general intelligence. Chollet himself said reaching human-level would “take many years.”
Then December 2024 happened. OpenAI’s o3-preview hit 87.5% in high-compute mode. First time any AI system crossed the human threshold of 85%. The model needed 1,024 attempts per task, writing roughly 137 pages of reasoning per attempt. Cost estimates ranged from $3,000 to $30,000 per task.
Eleven months later, GPT-5.2 Pro hits 90.5% at $11.64 per task.
The math on that cost collapse tells you everything. At $30,000 per task, you’d need to pay a human $6,000/hour to match the economics. At $11.64, a Mechanical Turk worker at $5/task is now more expensive than frontier AI reasoning. We crossed the human-cost parity line sometime in the last few months and most people missed it.
Now zoom out to the competitive dynamics.
Three weeks ago, Google dropped Gemini 3. Topped the LMArena leaderboard at 1501 Elo. Set records on Humanity’s Last Exam. Sam Altman publicly praised it. OpenAI declared “code red” internally, shelved projects like ad integrations, and fast-tracked GPT-5.2’s release from later this month to today.
This is the first model launch in OpenAI’s history that was explicitly a response to a competitor. The Verge reported employees asked to delay the release for more polish. Leadership overruled them. The directive was to reclaim the performance lead now.
And on ARC-AGI, they did. GPT-5.2 Pro at 90.5% edges out everything else on the board. But the real competition isn’t on accuracy anymore. Look at the cost-per-task column. The battle has shifted from “who can solve it” to “who can solve it cheaply.”
The efficiency gains aren’t slowing down. They’re compounding. Every major lab is now competing on the same benchmark, which means the collective R&D spend attacking this problem is in the billions. The 2025 ARC Prize Grand Prize ($700,000 for 85% on the private eval with efficiency constraints) is almost certainly getting claimed.
What happens after ARC-AGI-1 falls completely?
Chollet already released ARC-AGI-2 in March 2025, specifically designed to be harder for reasoning systems. Humans still hit nearly 100%. Current frontier models manage 10-45%. The gap between human and AI performance on even the harder benchmark is now a cost optimization problem, not a fundamental capability barrier.
If you’re building products in 2025 and assuming AI reasoning is expensive, you’re building for a world that no longer exists.
The benchmark that was supposed to prove AI couldn’t generalize just became another line item on a pricing page. 390X efficiency improvement in one year.
TOON (Token-Oriented Object Notation) is out for some days now and it aims to make communication with LLMs more accurate and token-efficient.
The TOON topic is now one of the hottest news on the LLM market and it might actually matter.
𝗪𝗵𝘆 𝗜 𝘁𝗵𝗶𝗻𝗸 𝘀𝗼:
I was initially hesitant to cover this, potentially being another hype to quickly fade, but:
✅ The format has been shown to increase the accuracy of models while decreasing the token count. I was not sure if there were any accuracy retention studies made, it seems there were.
✅ Token efficiency is extremely important when working with Agentic Systems that require a lot of structured context inside of their reasoning chains. And we are moving towards a post-PoC world where there is a lot of emphasis placed on optimisation of the workflows.
𝗔 𝘀𝗵𝗼𝗿𝘁 𝘀𝘂𝗺𝗺𝗮𝗿𝘆:
- Token-efficient: typically 30-60% fewer tokens on large uniform arrays vs formatted JSON.
- LLM-friendly guardrails: explicit lengths and fields enable validation.
- Minimal syntax: removes redundant punctuation (braces, brackets, most quotes).
- Indentation-based structure: like YAML, uses whitespace instead of braces.
- Tabular arrays: declare keys once, stream data as rows.
An example:
𝘑𝘚𝘖𝘕 𝘧𝘰𝘳𝘮𝘢𝘵:
"shopping_cart": [
{ "id": "GDKVEG984", "name": "iPhone 15 Pro Max", "quantity": 2, "price": 1499.99, "category": "Electronics" },
{ "id": "GDKVEG985", "name": "Samsung Galaxy S24 Ultra", "quantity": 1, "price": 1299.99, "category": "Electronics" },
{ "id": "GDKVEG986", "name": "Apple Watch Series 9", "quantity": 1, "price": 199.99, "category": "Electronics" },
{ "id": "GDKVEG987", "name": "MacBook Pro 16-inch", "quantity": 1, "price": 2499.99, "category": "Electronics" }
]
}
𝘞𝘩𝘦𝘯 𝘦𝘯𝘤𝘰𝘥𝘦𝘥 𝘪𝘯𝘵𝘰 𝘛𝘖𝘖𝘕 𝘧𝘰𝘳𝘮𝘢𝘵:
shopping_cart: items[4]{id,name,quantity,price,category}:
GDKVEG984,iPhone 15 Pro Max,2,1499.99,Electronics
GDKVEG985,Samsung Galaxy S24 Ultra,1,1299.99,Electronics
GDKVEG986,Apple Watch Series 9,1,199.99,Electronics
GDKVEG987,MacBook Pro 16-inch,1,2499.99,Electronics
𝗥𝗲𝘀𝘂𝗹𝘁:
✅ 43% savings in token amount.
✅ Directly translates to 43% savings in token cost for this LLM input.
❗️ Be sure to know when NOT to use the format (and always test it for your application specifically):
- Deeply nested or non-uniform structures.
- Semi-uniform arrays.
- Pure tabular data.
ℹ️ I will be testing it in the upcoming weeks.
Let me know if you have already tested TOON and what are your takeaways! 👇
#LLM #AI #MachineLearning
1. Drink water early in the morning. Do it for your kidneys.
2. Eat at least two eggs every day. Do it for your brain and hormones.
3. Practice Kegel exercises and take a walk. Do it for your heart.
4. Drink a shot of ginger every morning. Do it to boost your immune system.
5. Get morning sunlight. Do it for your skin.
6. When you sit, sit straight. When you stand, stand straight. Do it for your spine.
7. Eat raw garlic before bedtime. Do it to boost your testosterone.
8. Eat raw onions or add them to your meals. Do it to help prevent cancer.
Repost for others to learn!
En #CaixaForum València inauguramos “Dinosaurios de la Patagonia” con la colaboración de @mefpatagonia. Una experiencia que te sumerge en las historias de la tierra y los seres que la habitaron. #CaixaForumDinosaurios
This is the great and terrifying paradox of the AI era, and "model collapse" is the perfect term for it.
But this process doesn't just destroy value; it also creates it. As the cost of generating generic, average content trends toward zero, the economic value of two things skyrockets:
Verifiable, authentic, human-generated data for training the next generation of models.
Genuine, niche, human expertise that stands out from the sea of synthetic sameness.
The internet isn't dying. It's undergoing a great sorting. The future is about curation and finding the verified human signal in the AI noise. Authenticity is about to become our most valuable asset.
Oxford researchers just confirmed what we feared:
The internet as we knew it is dying.
AI content went from ~5% in 2020 to 48% by May 2025. Projections say 90%+ by next year.
Why? AI articles cost <$0.01. Human writers cost $10-100.
But the real crisis is model collapse. When AI trains on AI-generated content, quality degrades like photocopying a photocopy. Rare ideas disappear. Everything converges to generic sameness.
It's recursive. Today's AI slop becomes tomorrow's training data, producing worse output, which becomes training data again.