I’ve been using Codex for my day-to-day work as well.
So far, the app works much better, and the integrations make it easier to work with Slack and documents. I also feel that its writing has improved significantly.
Whenever I generate something with Claude, Codex finds an issue with it, and Claude always agrees. That rarely happens the other way around. Codex’s output is much cleaner.
Si pagas el plan de 200 de anthropic como yo hasta hace 5 minutos, te recomiendo hacer downgrade a la de 5x y comprarte la de 100€ de codex.
Desde hace mes y medio combino ambos y ha sido mi mejor decisión.
Codex para trabajar desde la app, revisar lo que hace claude y integrarlo con todo tu trabajo.
Claude para diseñar, y proyectos más profundos.
Not going back
Al final todo lo trabajo en local @ahrrruiz! Por lo tanto ambos tienen acceso a la misma carpeta.
Otra cosa que hago es que como uso claude en el terminal copio todo el rollo del terminal y se lo pego tal cuál a codex.
Y la última es que todo el output de uno y otro va etiquetado siempre. _donwclaude o _donwcodex y siempre se vuelca en una carpeta de output por mes.
Hoy vas a leer a todo el mundo que Anthropic ha lanzado el modelo que lo cambia todo.
400€ después te digo que es inutilizable.
Como buen friki del tema, tenía que probarlo y sí, básicamente es inutilizable para tareas reales.
Le he puesto a hacer mientras dormía una tarea que ya hice en su momento con 4.6 y 4.7.
Procesar las más de 3k transcripciones de grabaciones que tengo de mentorías y sesiones con clientes para mapear frameworks de trabajo, metodología y demás.
Me he despertado 6 horas después sin resultado (el token limit de la sesión había saltado) y con 150€ de extra tokens gastados.
Durante la noche, @Anthropic ha reseteado los tokens del modelo y ya va por el 40 por ciento (la captura desde hace un rato).
A la hora de escribir esto sigo sin output y consumiendo tokens.
Quiero ver el output que genera, pero el coste lo hace impracticable; es una tarea que en 4.6 max hice sin problema con un resultado decente.
Sabiendo lo que fallan las IAs, usar este modelo es hacer un salto de fe de 600€ cada vez y creo que hace de este un modelo para Twitter y LinkedIn, pero no para la vida real.
Voy a completar la tarea cueste lo que cueste; a ver si me hace cambiar de opinión.
BREAKING:
Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world.
We've been testing it internally @every for the last week or so across coding, writing, marketing, editing, and more—here's our vibe check:
- It broke our benchmarks. Fable scored a 91/100 on our Senior Engineer benchmark—this is human senior engineer level. The previous high score was Opus 4.8 at 63. GPT-5.5 is a 62.
- It's a one-shot wonder. You can set it and forget for hours or overnight on huge coding tasks, and come back to completed work. It cleared entire production bug backlogs, built a playable 3D, and even made a 2-minute animated film—all one-shot.
- Taste and attention to detail. In coding and knowledge work tasks, it has much better taste and attention to detail than we've ever seen. It gets subtle things right, adds little features you might not have thought of, and generally understands the assignment in ways that surprised us.
- Great use of context. We set it loose analyzing customer feedback surveys and our website data and it came back with a crisp, clean report that identified a. our biggest problem and b. a concrete testable solution—and then we sent it off to build that.
- It's best for power users. If you're already used to orchestrating multiple agents in your work, this model can do things that you've never seen before. If you're a knowledge worker or vibe coder with a more basic setup, you're not going to notice a huge difference—in fact, it probably isn't the right model for you.
- It's very slow, token-hungry. Using this thing for regular knowledge work is like squashing an ant with a rocket launcher. It also routinely uses 500k to 1M tokens on tasks. That's why it's best for your heaviest jobs—but not as good for tasks like collaborative writing.
- It's expensive. It's about twice as expensive as Opus, and it's also incredibly token hungry—so expect it to be something you'll use sparingly unless your company pays for it.
Overall, I think of it like a warp drive for coding: It can get you across the galaxy in a few hours, when it used to take months or years. But it's not appropriate for getting around town—you need something faster, cheaper, and more maneuverable.
The ceiling is extraordinarily high on this model though. Even our most advanced testers like @kieranklaassen felt like they were only scratching the surface of it.
Want our full vibe check with all of our testing and benchmarks? Read it on @every: https://t.co/MgJLZszJUB
Today, we're introducing Claude Fable 5 and Mythos 5, two configurations of our next major language model.
I'd normally highlight the numbers: It's SOTA on nearly all benchmarks. I want to talk about something else, because with Fable 5 out in the world, I think a third era quietly started today.
I lead Claude Code & Cowork on the desktop, so I think a lot about how people use AI to get work done. I believe we're about to see a major shift, moving from giving AI tasks to giving it responsibilities.
I use Claude Code as my main driver and reach for Codex for verification. Same projects, different surfaces: Claude Code lives in my terminal with the repo, Codex I use in the app.
The split that emerged: Claude drafts, the other model audits facts that have a ground truth.
Claude Code was my only AI system until ~2 weeks ago. The entire repo is built for Claude.
Codex got its AGENTS.md 5 days ago, with one section. It plays in Claude's house.
Last week I drafted an administrative appeal in Spanish. Claude wrote 3 Spanish Supreme Court citations (case numbers, years, paragraphs).
All 3 were invented. ChatGPT-5.5 caught them across 3 review rounds before I filed.
Same week, different domain: Codex caught 7 specific errors in a Meta Ads causal analysis Claude wrote.
"0 creative changes in window" was actually 2 update_ad_creative rows in the raw CSV.
A YoY table wasn't supported by any JSON on disk. It validated the central thesis but corrected the supporting numbers.
Main frustration with Claude: even when the whole context is built for it, it commits to a causal narrative or invents specific citations when the data only supports correlation, or "I don't know yet."
If Claude pushed back on its own prose with "is this verifiable in the raw source" before drafting, my reach-for-other-model rate would drop a lot.
(All of the above with xhigh or max effort)
Llevo unos días combinando codex con code y me está empezando a impresionar que cada vez que codex revisa algo y se lo paso a code de vuelta, la respuesta siempre está siendo...
No te estás dando cuenta, pero estás bajando el nivel.
Hace unos días me di cuenta de que la IA nos está haciendo bajar el nivel en general.
Y no porque el output que genera sea peor, sino porque te conformas antes con lo que te entrega.
Es decir:
. Landing A hecha con IA
. Landing B exactamente igual que la A hecha por un humano.
Por lo que estoy viendo, tanto a nivel interno como con clientes que la están usando mucho, el nivel y volumen de feedback sería mucho menor en la A que en la B.
Tu capacidad crítica se reduce cuando el output no lo genera un humano.
Seguro que existe un sesgo cognitivo que, si no está acuñado, habrá que acuñar.
No sé si es porque nos parece impresionante que un robot sea capaz de hacer eso, porque tarda menos que una persona o porque estás haciendo cosas que antes no podías hacer.
Pero la realidad es que eres más permisivo con la IA que con tu equipo y eso te hace bajar el nivel.
Llevo un par de meses notando esa tendencia que se agrava cada vez más, porque cada vez se pueden hacer más cosas.
Así que, como consejo, cada vez que revises algo que te ha hecho la IA, piensa...
¿Qué le diría a alguien de mi equipo si me hubiera entregado lo que me ha entregado la IA?
¿Sería más crítico/a?
Las one person billion dolar companies son unos titulares preciosos.
Pero nadie se ha dado cuenta de que añadir personas al negocio es la única opción para diluir el riesgo?
Qué pasa si el founder muere o se pone enfermo?
Aquí meto también empresas de 10 personas.
En un momento en el que la implementación es prácticamente fricción cero (todavía no estamos ahí) sigue teniendo sentido el modelo Hormozi de regala el conocimiento, vende la implementación?