Le puse memorias a mi agente: instrucciones, resúmenes, historiales.
Sigue sin aprender.
Puede explicarme perfectamente por qué falló ayer y cometer el mismo error hoy.
El problema no es cuánto guarda.
El cerebro esto lo resolvió durmiendo. https://t.co/lBVWrKRftO
Hace ya tiempo que le pedí a Grok que me hiciera una automatización para hacerme un reporte de novedades de IA cada día. Pero el tono era bastante negativo:
- Esto es malo
- Estafa...
Hoy fui a mejorar el reporte y encontré esto.
100% cierto, y más en código: el inglés es más conciso, y como los tokenizers están entrenados mayormente en inglés, el español consume 15-25% más tokens para decir lo mismo.
De todas maneras si no manejás bien el inglés, prompteá en español. Claridad > ahorro de tokens.
No es el fin del mundo promptear en español.
Lo que es hablar sin saber eh.
Jung es literalmente el de la sincronicidad, la alquimia y los horóscopos. Una hipótesis más incomprobable que la otra.
Increíble seguir hablando de esos dos en Big 2026 cómo si fueran ciencia
@tebayoso El de OpenAI de 20 es espectacular. Tengo Claude Max y el límite de 5 horas de Claude me queda súper corto. Codex me termina salvando las papás y encima los límites son altísimos, pocas veces me quedé sin usage. Tengo pendiente probar grok más a fondo
@lendersacc No creo que se degrade mucho el output, para texto me imagino que será como tiene Gemini con SynthId que elige ciertas palabras en lugar de otras. Para código casi no hay forma más allá de algunas elecciones de wording, así que dudo que ahí funcione la detección
@maxifirtman Me da mucha duda cómo van a implementar esto, si lo hacen con algo como SynthID hay varias formas de hacerle bypass y en código directamente no funciona. Ya vamos a ver igual varias páginas para evitar los watermarks
Hermoso.
Décadas de matemáticos para llegar a 41.6% y un modelo lo subió a 67.2% casi de casualidad, intentando otra cosa, mientras recibía inputs tipo “keep going” y “believe in yourself”.
Imaginate lo que viene para el resto.
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
Esto es adonde va todo, no tiene sentido leer el código y tampoco van a tener sentido los frameworks, herramientas para dev, etc.
Si seguís leyendo código y viendo manualmente línea por línea te convertís en un bottleneck gigante.
I'm officially done reading AI-generated code.
It's been two weeks since I looked at any of it.
I think the IDE is officially on its way to the graveyard. The job is no longer about "writing code," so we need new tools that better reflect this new reality.
While reviewing the code, I realized my only complaints were stylistic, and I wasn't finding any obvious bugs anymore.
The more code I generated, the harder it became to keep track of every line. I found that my time is better spent designing ways to verify that the overall system works than looking at the code.
State-of-the-art coding agents are better at writing code than I'd ever be, and I'm going to stop pretending otherwise.
I still think these coding agents can't go too far without an experienced human guiding them, but we're past the point where we need to check every line of code.
Introducing Kitesurf: a browser built for agents, running entirely on Cloudflare Workers.
Chromium is too heavy to hand every agent one. Kitesurf is written in Rust, uses 3-7x less CPU and memory, and spins up per request.
Free in beta: https://t.co/Vm8kVCgSkG