Aquela hora que você vai dormir deixando uma tarefa longa pro claudinho executar a noite. Acorda e vê que ele não fez porr nenhuma (mesmo eu estando em yolo mode). Porr claudinho ....
@felipepoloruiz Haz que todo el equipo se lea esta serie de artículos y los elevas al siguiente nivel en una hora #ShamelessSelfPromotion https://t.co/T5XNbFSxHp
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
How I work with AI today:
* Codex and Claude Code, almost never the chat.
* Folders connected to Git repos. mirrored on a server (VPS)
* I talk to them (voice) a lot at the beginning to refine a plan. No more prompt engineering.
* The server lets them run things while my laptop is offline.
* For each task, I ask for 10 alternatives, pick the best one, and ask them to remember my feedback.
* I make Fable 5 and GPT 5.6 review each other's work.
* I give them access to everything, no limits.
I don't read code anymore. There is no point to it. But I build graphs of the whole system and update them often so that I can have a mental model of the product in my mind.
Si tuviese que elegir 1 sola acción para restaurar el estres y empezar a recuperar la calma mental...
No sería dejar el móvil 1 hora antes de acostarme. No sería amanecer sin móvil. No sería dejar las RRSS.
La clave real para empezar bien es:
Everything that matters most in life will sometimes bore you.
Your marriage will bore you.
Your career will bore you.
Your kids will bore you.
But boredom isn’t a sign that you chose wrong. It’s a signal that you have chosen to make something more important than your own momentary pleasure. It’s the hidden fee that comes with purpose and meaning.
¿Pensando en devolverte a Venezuela?
Regresar a Venezuela no es volver a casa. Es volver a un lugar que tiene el mismo nombre que tu casa, la misma geografía, los mismos olores, pero que ya no es lo que dejaste. Y tú tampoco eres quien eras cuando te fuiste. Dos cosas que cambiaron tratando de encajar como si el tiempo no hubiera pasado. Eso no es un regreso, es una colisión.
El problema no es extrañar, el problema es confundir el país que llevas en la memoria con el país que existe hoy. Extrañas la arepa de tu mamá, el calor de tu barrio, la carcajada fácil con los tuyos, y eso es completamente legítimo. Pero ninguna de esas cosas resuelve la inflación, la inseguridad, la ausencia de futuro, la sensación de que por más que trabajes el piso se mueve solo.
Regresar por nostalgia es querer curar el hambre mirando una foto de comida.
Y lo más duro de todo es que quien regresa no solo pierde lo que construyó afuera.
Pierde también la versión de sí mismo que estaba emergiendo, esa persona más fuerte, más adaptable, más libre, que estaba aprendiendo a existir sin red. Venezuela puede esperar a que las cosas cambien. Tú no puedes esperar a que cambies de nuevo.
Obsidian, Knowledge Base Manager, es formalmente LLM native:
- Está basado en una jerarquía de directorios con archivos Markdow
- Todo lo que puedes hacer el la UI está soportado en el CLI
Y pronto tendrá soporte headless.
Y es Open Source y gratuito.
Happy 2026! Will this be the year we finally achieve AGI? I’d like to propose a new version of the Turing Test, which I’ll call the Turing-AGI Test, to see if we’ve achieved this. I’ll explain in a moment why having a new test is important.
The public thinks achieving AGI means computers will be as intelligent as people and be able to do most or all knowledge work. I’d like to propose a new test. The test subject — either a computer or a skilled professional human — is given access to a computer that has internet access and software such as a web browser and Zoom. The judge will design a multi-day experience for the test subject, mediated through the computer, to carry out work tasks. For example, an experience might consist of a period of training (say, as a call center operator), followed by being asked to carry out the task (taking calls), with ongoing feedback. This mirrors what a remote worker with a fully working computer (but no webcam) might be expected to do.
A computer passes the Turing-AGI Test if it can carry out the work task as well as a skilled human.
Most members of the public likely believe a real AGI system will pass this test. Surely, if computers are as intelligent as humans, they should be able to perform work tasks as well as a human one might hire. Thus, the Turing-AGI Test aligns with the popular notion of what AGI means.
Here’s why we need a new test: “AGI” has turned into a term of hype rather than a term with a precise meaning. A reasonable definition of AGI is AI that can do any intellectual task that a human can. When businesses hype up that they might achieve AGI within a few quarters, they usually try to justify these statements by setting a much lower bar. This mismatch in definitions is harmful because it makes people think AI is becoming more powerful than it actually is. I’m seeing this mislead everyone from high-school students (who avoid certain fields of study because they think it’s pointless with AGI’s imminent arrival) to CEOs (who are deciding what projects to invest in, sometimes assuming AI will be more capable in 1-2 years than any likely reality).
The original Turing Test, which required a computer to fool a human judge, via text chat, into being unable to distinguish it from a human, has been insufficient to indicate human-level intelligence. The Loebner Prize competition actually ran the Turing Test and found that being able to simulate human typing errors — perhaps even more than actually demonstrating intelligence — was needed to fool judges. A main goal of AI development today is to build systems that can do economically useful work, not fool judges. Thus a modified test that measures ability to do work would be more useful than a test that measures the ability to fool humans.
For almost all AI benchmarks today (such as GPQA, AIME, SWE-bench, etc.), a test set is determined in advance. This means AI teams end up at least indirectly tuning their models to the published test sets. Further, any fixed test set measures only one narrow sliver of intelligence. In contrast, in the Turing Test, judges are free to ask any question to probe the model as they please. This lets a judge test how “general” the knowledge of the computer or human really is. Similarly, in the Turing-AGI Test, the judge can design any experience — which is not revealed in advance to the AI (or human subject) being tested. This is a better way to measure generality of AI than a predetermined test set.
AI is on an amazing trajectory of progress. In previous decades, overhyped expectations led to AI winters, when disappointment about AI capabilities caused reductions in interest and funding, which picked up again when the field made more progress. One of the few things that could get in the way of AI’s tremendous momentum is unrealistic hype that creates an investment bubble, risking disappointment and a collapse of interest. To avoid this, we need to recalibrate society’s expectations on AI. A test will help.
If we run a Turing-AGI Test competition and every AI system falls short, that will be a good thing! By defusing hype around AGI and reducing the chance of a bubble, we will create a more reliable path to continued investment in AI. This will let us keep on driving forward real technological progress and building valuable applications — even ones that fall well short of AGI. And if this test sets a clear target that teams can aim toward to claim the mantle of achieving AGI, that would be wonderful, too. And we can be confident that if a company passes this test, they will have created more than just a marketing release — it will be something incredibly valuable.
[Original text: https://t.co/mGAmoOGga7 ]
I can't believe Anthropic got this far
No image models, no audio models, no huge context windows, nothing, just models focused on being excellent at code
The Claude website lacks many features and the rate-limits are bad, yet it's still a strong competitor for ChatGPT/Gemini
𝗜 𝗱𝗼𝗻'𝘁 𝘄𝗿𝗶𝘁𝗲 𝗰𝗼𝗱𝗲 𝗮𝘁 𝗚𝗼𝗼𝗴𝗹𝗲 𝗮𝗻𝘆𝗺𝗼𝗿𝗲. 𝗜 𝗿𝗲𝘃𝗶𝗲𝘄 𝗶𝘁.
𝙄𝙛 𝙄 𝙡𝙤𝙤𝙠 𝙖𝙩 𝙢𝙮 𝙙𝙖𝙞𝙡𝙮 𝙘𝙤𝙢𝙢𝙞𝙩𝙨, 70-80% 𝙤𝙛 𝙩𝙝𝙚 𝙘𝙤𝙙𝙚 𝙞𝙨 𝙣𝙤𝙬 𝙬𝙧𝙞𝙩𝙩𝙚𝙣 𝙗𝙮 𝘼𝙄.
My role has fundamentally shifted.
• I don't type syntax; I prompt logic.
• I don't hunt for bugs; I review AI's suggestions.
• I don't read legacy code; I ask AI to explain it.
A lot of engineers feel guilty about this. They feel like they are "cheating."
They aren't. They are evolving.
𝗜 𝗼𝗻𝗰𝗲 𝗮𝘀𝗸𝗲𝗱 𝗮 𝘀𝗲𝗻𝗶𝗼𝗿 𝗹𝗲𝗮𝗱𝗲𝗿 𝘁𝗵𝗲 𝗾���𝗲𝘀𝘁𝗶𝗼𝗻 𝗲𝘃𝗲𝗿𝘆𝗼𝗻𝗲 𝗶𝘀 𝗮𝗳𝗿𝗮𝗶𝗱 𝗼𝗳: "𝗪𝗶𝗹𝗹 𝗔𝗜 𝗿𝗲𝗽𝗹𝗮𝗰𝗲 𝘂𝘀?" 𝗛𝗶𝘀 𝗮𝗻𝘀𝘄𝗲𝗿 𝘀𝘁𝘂𝗰𝗸 𝘄𝗶𝘁𝗵 𝗺𝗲:
"𝘈𝘐 𝘪𝘴 𝘢 𝘮𝘶𝘭𝘵𝘪𝘱𝘭𝘪𝘦𝘳, 𝘯𝘰𝘵 𝘢 𝘳𝘦𝘱𝘭𝘢𝘤𝘦𝘮𝘦𝘯𝘵. 𝘐𝘧 𝘺𝘰𝘶 𝘶𝘴𝘦𝘥 𝘵𝘰 𝘥𝘰 1𝘹 𝘸𝘰𝘳𝘬 𝘪𝘯 𝘢 𝘸𝘦𝘦𝘬, 𝘵𝘩𝘦 𝘦𝘹𝘱𝘦𝘤𝘵𝘢𝘵𝘪𝘰𝘯 𝘪𝘴 𝘯𝘰𝘸 4𝘹 𝘸𝘰𝘳𝘬 𝘪𝘯 𝘵𝘩𝘦 𝘴𝘢𝘮𝘦 𝘸𝘦𝘦𝘬. 𝘕𝘰 𝘤𝘰𝘮𝘱𝘢𝘯𝘺 𝘸𝘢𝘯𝘵𝘴 𝘵𝘰 𝘮𝘰𝘷𝘦 𝘣𝘢𝘤𝘬𝘸𝘢𝘳𝘥."
The bar for "productivity" has moved.
If you refuse to use AI because you are a "purist," you aren't noble. You are just slow.
𝗔𝗜 𝘄𝗶𝗹𝗹 𝗻𝗼𝘁 𝗿𝗲𝗽𝗹𝗮𝗰𝗲 𝘆𝗼𝘂. 𝗕𝘂𝘁 𝗮𝗻 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿 𝗱𝗼𝗶𝗻𝗴 4𝘅 𝘁𝗵𝗲 𝘄𝗼𝗿𝗸 𝘂𝘀𝗶𝗻𝗴 𝗔𝗜... 𝗱𝗲𝗳𝗶𝗻𝗶𝘁𝗲𝗹𝘆 𝘄𝗶𝗹𝗹.
I'm not joking and this isn't funny. We have been trying to build distributed agent orchestrators at Google since last year. There are various options, not everyone is aligned... I gave Claude Code a description of the problem, it generated what we built last year in an hour.
We are SO excited about the response to Infographics and Slide Decks! Due to the overwhelming demand, we're experiencing some capacity constraints and have temporarily rolled back access to these features for Free users and instituted additional limits on generations for Pro users, however we plan on bringing everything back to normal as soon as we can!
Thank you so much for your patience and understanding! 🙏