24 hours of Astra 6 ... and i think it sucks for serious dev work - planning + coding. Very painful to work with. Wasting so much time watching it "think". Stopping after you give clear instructions. It wants to think more, so it stops. You come back to your screen expecting something remotely amazing, and nothing. From Light to Extra High, all the settings suck. What a let down. Back to 5.6 Sol.
@joshmanders same. I use codex to check the opus 5 plans before i execute in a 2nd opus 5 session, and codex keeps finding errors in every single plan. And the opus 5 coding does not inspire confidence. I feel opus 5 is very inconsistent.
@MaaSonder i feel the same way, Opus 4.6 > Opus 5. Too many simple mistakes. If i didn't use codex to check each plan before it runs, i'd have a mess on my hands.
A mulher que construiu o ChatGPT saiu da OpenAI, ficou em silêncio por um ano, e o que ela acabou de lançar pode mudar pra sempre como você usa IA no dia a dia.
Mira Murati não fundou mais um chatbot. Ela foi atrás do problema que nenhum lab quis resolver: toda IA que existe hoje funciona por turnos. Você digita, espera. O modelo responde, espera. É tentar resolver uma crise por e-mail quando você poderia estar na mesma sala que a pessoa.
O que a Thinking Machines lançou hoje acaba com isso.
O modelo ouve, vê, fala, pensa e age ao mesmo tempo. Não é um pipeline costurado de componentes. É o modelo em si que foi treinado do zero pra funcionar assim.
→ Latência de 0,40s por turno. O padrão da indústria é 1 a 2 segundos.
→ Micro-turnos de 200ms intercalando input e output sem parar
→ Faz busca, usa ferramentas e gera interface enquanto conversa com você
→ Percebe quando você hesita e intervém antes de você pedir
→ Tradução simultânea em tempo real com as duas partes falando
A equipe: Mira Murati como CEO (ex-CTO da OpenAI), Soumith Chintala como CTO (criador do PyTorch), e contratações recentes da Meta em percepção multimodal.
O ponto técnico que vale gravar: eles citam o "bitter lesson" do Rich Sutton. Interatividade construída por componente externo sempre vai perder pra interatividade nativa ao modelo. Escalar o modelo o torna mais inteligente e mais colaborativo ao mesmo tempo.
822 mil visualizações em 4 horas. a16z comentando. Brasil dormindo.
Toda IA que você usa hoje vai parecer e-mail dentro de dois anos. E quem largou na frente dessa corrida não foi OpenAI, Google nem Anthropic.
Foi a empresa da mulher que eles deixaram sair.
Yann LeCun was right the entire time. And generative AI might be a dead end.
For the last three years, the entire industry has been obsessed with building bigger LLMs. Trillions of parameters. Billions in compute.
The theory was simple: if you make the model big enough, it will eventually understand how the world works.
Yann LeCun said that was stupid.
He argued that generative AI is fundamentally inefficient.
When an AI predicts the next word, or generates the next pixel, it wastes massive amounts of compute on surface-level details.
It memorizes patterns instead of learning the actual physics of reality.
He proposed a different path: JEPA (Joint-Embedding Predictive Architecture).
Instead of forcing the AI to paint the world pixel by pixel, JEPA forces it to predict abstract concepts. It predicts what happens next in a compressed "thought space."
But for years, JEPA had a fatal flaw.
It suffered from "representation collapse."
Because the AI was allowed to simplify reality, it would cheat. It would simplify everything so much that a dog, a car, and a human all looked identical.
It learned nothing.
To fix it, engineers had to use insanely complex hacks, frozen encoders, and massive compute overheads.
Until today.
Researchers just dropped a paper called "LeWorldModel" (LeWM).
They completely solved the collapse problem.
They replaced the complex engineering hacks with a single, elegant mathematical regularizer.
It forces the AI's internal "thoughts" into a perfect Gaussian distribution.
The AI can no longer cheat. It is forced to understand the physical structure of reality to make its predictions.
The results completely rewrite the economics of AI.
LeWM didn't need a massive, centralized supercomputer.
It has just 15 million parameters.
It trains on a single, standard GPU in a few hours.
Yet it plans 48x faster than massive foundation world models. It intrinsically understands physics. It instantly detects impossible events.
We spent billions trying to force massive server farms to memorize the internet.
Now, a tiny model running locally on a single graphics card is actually learning how the real world works.
There's a physicist at Stanford named Safi Bahcall who modeled this exact principle and the math is wild.
He calls it "phase transitions in human networks." When you're stationary, your probability of a lucky event is limited to your existing surface area: the people you already know, the places you already go, the ideas you've already been exposed to. Your opportunity window is fixed.
When you move, your collision rate with new nodes in a network increases nonlinearly. Double your movement (new conversations, new cities, new projects) and your probability of a serendipitous encounter doesn't double. It roughly quadruples. Because each new node connects you to their entire network, not just to them.
Richard Wiseman ran a 10-year study at the University of Hertfordshire tracking self-described "lucky" and "unlucky" people. The single biggest differentiator wasn't IQ, education, or family money. Lucky people scored significantly higher on one trait: openness to experience. They talked to strangers more, varied their routines more, and said yes to invitations at nearly twice the rate.
The "unlucky" group followed the same routes, ate at the same restaurants, and talked to the same 5 people. Their networks were closed loops. No new inputs, no new collisions.
Luck isn't random. Luck is surface area. And surface area is a function of movement.
The lobster emoji is doing more work than most people realize. Lobsters grow by shedding their shell when it gets too tight. The growth requires a period of total vulnerability. No protection, no armor, soft body exposed to the ocean.
That's the cost of movement nobody posts about. You have to be uncomfortable first. The new shell only hardens after you've already moved.
This is big... Anthropic just announced a model so powerful they won't release it to the public out of fear over the damage it will cause 😨
Claude Mythos Preview found thousands of zero-day exploits in every major operating system and web browser...
The numbers are hard to believe:
> $50 to find a 27-year-old bug in OpenBSD, one of the most security-hardened operating systems ever built
> Under $1,000 to find AND build a fully working remote code execution exploit on FreeBSD that grants unauthenticated root access from anywhere on the internet
> Under $2,000 to chain together multiple Linux kernel vulnerabilities into a complete privilege escalation exploit
For context: these are the kinds of findings that previously required elite security researchers working for weeks.
Anthropic engineers with no formal security training asked Mythos to find exploits overnight. They woke up to working code the next morning.
The results were so impressive Anthropic assembled Apple, Google, Microsoft, Amazon, NVIDIA, and seven other organizations into Project Glasswing:
A $100M defensive coalition. They're not releasing this model publicly. Instead, they're racing to patch the world's infrastructure before models like this proliferate.
My friend Milla Jovovich and I spent months creating an AI memory system with Claude. It just posted a perfect score on the standard benchmark - beating every product in the space, free or paid.
It's called MemPalace, and it works nothing like anything else out there.
Instead of sending your data to a background agent in the cloud, it mines your conversations locally and organizes them into a palace - a structured architecture with wings, halls, and rooms that mirrors how human memory actually works.
Here is what that gets you:
→ Your AI knows who you are before you type a single word - family, projects, preferences, loaded in ~120 tokens
→ Palace architecture organizes memories by domain and type - not a flat list of facts, a navigable structure
→ Semantic search across months of conversations finds the answer in position 1 or 2
→ AAAK compression fits your entire life context into 120 tokens - 30x lossless compression any LLM reads natively
→ Contradiction detection catches wrong names, wrong pronouns, wrong ages before you ever see them
The benchmarks:
100% recall on LongMemEval — first perfect score ever recorded. 500/500 questions. Every question type at 100%.
92.9% on ConvoMem — more than 2x Mem0's score.
100% on LoCoMo — every multi-hop reasoning category, including temporal inference which stumps most systems.
No API key. No cloud. No subscription. One dependency. Runs on your machine. Your memories never leave.
MIT License. 100% Open Source.
https://t.co/KggwTqijmD
RIP "be specific in your prompts."
Everyone taught you the same thing. Add detail. Spell it out. Give the model every constraint it needs. The more specific, the better.
The research says otherwise.
A study titled "Same Task, More Tokens" found that LLM reasoning performance degrades significantly as prompt length increases and it kicks in well before you even get close to the model's context limit. We're talking degradation at around 3,000 tokens. Prompts most "advanced" users are already writing.
Here's what's actually happening under the hood.
When you stuff a prompt with constraints, formatting rules, edge cases, and redundant context, you're not giving the model more signal. You're giving it more noise. The attention mechanism has to spread across every token you added. The relevant parts get diluted. Reasoning collapses.
Researchers call it "Lost in the Middle" a documented phenomenon where models systematically fail to use information buried in long contexts, even when it's directly relevant. The model finds the beginning. The model finds the end. Everything in the middle becomes fog.
And the irony? The two prompting techniques everyone uses to handle complex tasks Chain-of-Thought and detailed instructions are the most vulnerable to this degradation when inputs get bloated.
So what actually works?
The shift the top AI engineers are making is from specification to direction.
Instead of listing 12 constraints, you define the outcome and give the model space to reason. Instead of "write a 300-word formal summary structured around three themes with a conclusion that references the intro," you write "synthesize this into a tight executive summary."
The model with room to think beats the model in a cage every time.
This is exactly why Anthropic moved its own internal guidance away from "prompt engineering" toward "context engineering." The game isn't cramming instructions. It's delivering the minimum high-signal context that unlocks the reasoning you want then getting out of the way.
Shorter prompts aren't lazy.
They're the sign you actually understand how the model thinks.
@ProductFaculty I don't think this works if you're a SaaS company that builds and ships commercial software. The PRD is the contract between dev, pm and pmm (marketing) - it helps define the roadmap and keeps everyone aligned on the march to release. Automate the PRD, not eliminate it.