@rohanpaul_ai LeCun's right that world models beat next-token prediction for robotics, and he's still the loudest guy in Silicon Valley shipping nothing.
@itsTarH The part nobody models is that a gas turbine takes three years to ship, so power availability now decides how fast a data center can open at all.
@cyrilXBT Self-hosted agents with permissions you can actually tighten is the boring part that makes this real, everyone else keeps shipping black boxes.
🧠 Claude Opus 5.5 uses 99% fewer em dashes than Opus 5.
That doesn't seem significant but , in my opinion they are one of the biggest telling clue in a text .9
Graphite ran a fingerprint audit across 9,974 topics and two weeks of versioning took one of the most public AI tells to almost nothing. Em dashes were a free, model specific fingerprint. That artifact is now 1% of what it was.
The tell layer never collapsed though. 2,548 tells in 5.5 versus 2,666 in Opus 5. Four percent fewer. The fingerprints just churn into new ones, so every detector has to be rebuilt on every release.
OpenAI went the other way. GPT-6 Astra sits 8% FURTHER from human word distribution than GPT-5.6.
Which means the detection economics run backwards: detectors pay recurring retraining cost and always lag, while the lab pays nothing and can revoke the signal for free.
If passing a detector is cheap, was it ever measuring anything?
🔬 Your chain-of-thought monitor catches the model 94% of the time. Add a normal system prompt and it catches almost nothing.
A MATS paper (Duzan and Stickland, arXiv 2608.04735) ran the same benchmark two ways.
The old way: tell the model explicitly to hide a side task. Monitor catches the leak in the scratchpad 60 to 94% of the time, across seven frontier reasoning models and four task formats.
The way production actually works. Nobody instructs anything. You just slip an irrelevant detail about a candidate into a resume, and the rating drifts.
Same bias. Detection fell 41 to 46 points in two of the four formats.
Then they added a realistic system prompt, exactly the kind you would ship to reduce off-topic bias. Implicit detection fell to as low as 5%. The influence stayed at full strength.
So what was the monitor reading? The vocabulary of concealment. Hide-this-then-do-that leaves a lexical fingerprint. A bias absorbed from context leaves nothing, because the model never decided to conceal anything. The trace is honest. It just faithfully records a conclusion it reached with no visible switch.
Nobody lied about those monitorability numbers. They were measured against loud concealment, and the quiet version is what runs in your hiring pipeline.
@rohanpaul_ai A 44-point swing from one instruction line means the alignment gap is way wider than leaderboards admit, and I'd bet every eval suite quietly drops the cheaters now.
@DimaZeniuk Orbital compute only makes sense once launch costs drop another order of magnitude, and SpaceX is the only operator positioned to move that cost curve itself.
Here is a quick breakdown of the massive updates that dropped this week
The AI landscape is moving so fast right now it's almost impossible to keep up
🔥 Gemini 4 Argon
• Massive benchmark leaps pushing Google's flagship reasoning capabilities to new heights.
• Deeper integration inside @Google Gemini Notebook, making long-context analysis and multi-modal workflows smoother.
🟢 @OpenAI DevDay Drop
• ChatGPT Dots: A whole new interaction layer designed to streamline continuous context and micro-tasks inside the chat UI.
• ChatGPT Space: Expanded workspace features built for deep collaborative projects.
• GPT-6.1 Sol: Next-generation model efficiency built for rapid reasoning and lower latency.
🟣 Claude Sonnet 5.5
• Anthropic doubles down on developer tooling and coding precision.
• Upgraded capabilities for complex multi-step execution and automated agent workflows.
🌐 Meta & Manus
• Meta Muse: New generative capabilities pushing multimodal creative limits.
• Manus Q: Autonomous agent capabilities continuing to evolve for seamless tool usage and real-world task execution.
The gap between standard LLM chat interfaces and fully-integrated autonomous work environments is closing fast.
@omarsar0 We've all been tuning tokens per step when cost per task solved is the real number, which is exactly why the same policy flips Qwen to Devstral.