🔮 ¡MIS PREDICCIONES IA del 2026! 🔮
Un año más aquí os traigo 24 ideas de lo que creo podría pasar en el mundo de la Inteligencia Artificial durante este año
24 predicciones que podéis votar y apoyar una a una usando el botón 💖
¡Comparte el hilo! En 12 meses verificamos 😄👇
En una dimensión paralela Gary Marcus tuvo razón, el deep learning chocó contra un muro y lo más apasionante de la decáda de los 2020s fue el desarrollo de los NFTs y el metaverso 😮💨
Vivimos en la mejor línea temporal sin duda.
Como os comentaba en el vídeo, Artificial Analysis se había quedado desactualizado como evaluación, y es algo que se ha notado con la salida de GPT-6 Astra, cuando lo ha situado al mismo nivel que GPT-Sol 5.6
Por ello hoy han actualizado el índice y ahora si se observa un orden más coherente. Fable 5.1 sigue a la cabeza, eso sí.
Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming
Intelligence Index v4.2 changelog:
+ AA-Briefcase, our agentic knowledge work evaluation with a private test set
+ @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages
- GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated
… plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness
This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January.
We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users.
Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned!
Intelligence Index v4.2 changes in detail:
➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied.
➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5.
➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure.
Key results:
➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google
➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier
➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
Una vez se abren las puertas de GPT-6 Astra vamos a ver como esta red social se llena de ejemplos que demuestren su rendimiento.
Por ahora lo que leo es muy positivo. Va a ser un finde divertido :)
GPT-6 Astra made this ps5 controller in threejs, no model comes even close.
This is a huge step up from 5.6 Sol, and even from Fable 5.1 id say. Crazy model.
@PabloSeitler Nice! Viendo las curvas de eficacia/coste te diría de no ponerlo a max effort, creo que en high o medium te desempeñará igual de bien y sin gastar tanto.
We are progressing through the rollout of Astra. Pro and Business subscriptions get it first, some of you should start seeing it across ChatGPT Work and Codex.
And then we will proceed with rollout to all of Plus as fast as we can.
En la generación de documentos muestran este ejemplo que no termina de entusiamarme. Se supone que muestra cómo GPT-6 se ajusta al estilo dado como referencia (a la izquierda) para hacer un .ppt como el de la derecha.
Sigo viendo detalles finos de diseño pasados por alto, como fuentes distintas, tamaños, esquinas no redondeadas, etc. que me hacen pensar que todavía queda margen de mejora para que aquí la IA haga un trabajo sobresaliente de primeras.
🔴 ¡¡OPENAI ANUNCIA GPT-6!!
El nuevo modelo GPT-6 Astra ya está aquí, con un salto en capacidades MUY sorprendente!
Os iré desglosando y analizando todos los detalles en este hilo 👇🧵
¡deja tu RT para apoyar!
Por no confundir, un matiz a tener en cuenta es que el uso de los looped transformers no viene a sustiuir irremediablemente el chain-of-thoughts por un pensamiento oculto.
Sí es cierto que añade más procesamiento interno por parte del modelo en el espacio latente, pero igual que si se hubiera decidido usar un tranformer convencional con más bloques de procesamiento (aquí con pesos compartidos), y eso en cualquier caso no quita que el CoT siga produciéndose.
El problema está en que a mayor procesamiento interno, ya sea por mayor profundidad, o escala del modelo o lo que sea, el modelo empieza a tener más control sobre sus propias CoT, que es lo que está empezando a generar propblemas de monitorización según han reportado en la system card de OpenAI.
Pero me parece un problema que probablemente iba a aparecer tarde o temprano al continuar con el escalado de los modelos, con looped transformers o sin ellos.
@palosdediego_ Sip, ChatGPT a finales del '22. Pero antes ya teníamos a GPT-3 y como modelo finetuneado para tareas de programación a Codex. Eramos pocos los que estábamos atentos a lo que se venía por aquellas fechas :)
¿Recordáis el nuevo Codex de OpenAI? 🤔
Pues ya tengo acceso y ya estoy haciendo pruebas locas con él. El comentario lo he escrito yo y el código lo ha generado automáticamente el sistema :)
🔴 ¡¡OPENAI ANUNCIA GPT-6!!
El nuevo modelo GPT-6 Astra ya está aquí, con un salto en capacidades MUY sorprendente!
Os iré desglosando y analizando todos los detalles en este hilo 👇🧵
¡deja tu RT para apoyar!
Oh wow, OpenAI quieren darnos GPT-6 Astra rapidito. Tanto que dicen que por cada día que se retrasen nos van a dar un reset acumulable desde el día de hoy.
We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can.
First one will land in ~ 3 hours. There is still time to create your account if you don't have one.