The TEDx talk I did is now available. It’s a summary of my view of what the key challenges of humanity are over the next 100 years. "Quo vadis" / Where are you going? | Esat Sibay | TEDxTUDarmstadt https://t.co/kjGHMD0JZw via @YouTube
OpenAI is buying Mac Mini's by the tens of thousands and the interesting part is what this says about where the bottleneck has moved.
The scarce input for this generation of models isn't only GPUs, it's realistic places for agents to practice. Labs already pay heavily for those places. Anthropic leadership has reportedly discussed spending over $1 billion on RL environments in the next year, and vendors like Surge (about $1.2B in revenue from the big labs) and Mercor (valued at $10B) are racing to build task environments (gyms) for them.
Why computer use needs real computers:
Both OpenAI and Anthropic are trying to build models that operate software the way a person does. Look at the screen, click, type, check the result. The training method is reinforcement learning at scale roughly following this process:
1. spin up thousands of real desktop sessions
2. hand the agent tasks like editing and testing code or organizing email
3. score the outcome
4. update the model on what worked
Practitioners describe single runs sustaining this loop for days or weeks, and most of that time goes to the agents working through tasks, not to updating the model weights.
The reason to focus on Mac is that macOS accounts for roughly a fifth of the world's desktops, second only to Windows, and overrepresented among the developers these agents are supposed to help. The only way to get a model that is expert at using a Mac is to have it use a Mac, millions of times. All that practice is showing up on the scoreboard between @OpenAI and @AnthropicAI.
Who's ahead right now in the Computer Use race:
The main eval is OSWorld-Verified and in this test, an agent gets 369 real desktop tasks and passes only if the machine ends up in the right state. Humans scored about 72% when the benchmark launched. The top models are now past that.
- Claude Fable 5 (Anthropic's flagship): 85% (https://t.co/8Yd0qGUozZ)
- GPT-5.5 (OpenAI's best published on this test): 78.7% (https://t.co/12ipJY4hDl)
NOTE: OpenAI's newest flagship, GPT-5.6 Sol, reports on a different variant: 62.6% on OSWorld 2.0, a newer and harder task set where it's the state of the art (https://t.co/xdjfNwcufk). These are different tests, so the numbers don't compare directly. The honest summary is Anthropic leads the older benchmark, OpenAI leads the newer one.
Why Mac minis and not macOS Virtual Machines in the cloud:
This part has to do with licensing, not physics. Apple's macOS license only permits macOS to run on Apple hardware, so there is no legal way to spin up a macOS VM on an ordinary server fleet the way you would a thousand Linux or Windows machines.
AWS's "Mac instances", which Anthropic uses, are literally racks of physical Mac minis, billed with a 24-hour minimum allocation to comply with the Apple macOS Software License Agreement. So if you want macOS desktops at scale, you buy Macs or you rent real ones. A headless $899 mini drawing 60-120W is the cheapest legal container for a macOS session that exists.
The rent vs buy math:
Buying an M6 Mac mini lists at $899. Renting the closest AWS host, an M4 mini, runs about $1.23 an hour.
So left on 24/7 that's roughly $10,800 a year, or about $6,000 with a 3-year commitment. This means that break even for a single Mini running around the clock, happens in about a month compared to the rental prices.
So Anthropic's route costs 7-12x more per box per year. What it buys: no cash up front, no data center ops for an unusual hardware class, and a fleet that scales down when a run ends.
My guess is it also signals commitment level. You buy what you know you'll use for years and rent what you're still deciding about. Amazon being Anthropic's biggest cloud partner and investor makes AWS-first the natural posture anyway.
What actually happens on these machines:
1. The minis get racked headless and imaged with sandboxed macOS sessions. They are sandboxed because a half trained agent can theoretically click anything, which means practitioners need to keep them off the open web so they don't delete data or buy $5,000 items.
2. Each session runs one agent attempt: screenshot in, action out, for trajectories up to hundreds of steps.
3. An automated verifier checks the end state. Did the code pass its tests, did the email get filed correctly?
4. The trajectories flow back to a GPU cluster where the weights actually update, and the improved model redeploys to the fleet. The Macs mainly generate the training data, and the gradient math almost certainly stays on Nvidia GPUs.
Procurement:
Compared to acquiring GPUs, acquiring Mac Minis is simple. Data center GPU capacity is allocated years in advance. Mac minis sit in retail channels you can buy from this week. The catch is that consumers share those channels.
High memory Mac configs have been out of stock for months and Apple pulled the 512GB Studio from sale at one point, while Mac revenue jumped 29% year over year to $10.3B. Apple even shipped its M6 mini refresh this month instead of the usual fall window.
My read:
GPUs became the commodity layer for training compute, and Mac minis are starting to become part of a commodity layer for something newer, generating agent experience. A $899 desktop is the cheapest legal unit of real macOS practice, and the labs are treating it as a component, not a computer.
HubX, a startup developing AI apps, has become Turkey’s first AI-related unicorn, reaching a valuation of $1.2 billion and the country’s 8th tech startup unicorn.
I believe Turkiye will emerge as one of the major winners of the AI age. The country has a remarkable number of high-agency and entrepreneurial people who rapidly embrace new technologies and from my recent observations, they are also now embracing AI with extraordinary speed. The Turkish government is also increasingly supporting and investing in the AI revolution.
The countries, companies, and individuals that go all in on AI will prosper greatly in the age of AI.
An important articles that articulates something I have been thinking for awhile really well: we need different models for Different tasks and a “model orchestrator” that “continuously moves each workload toward the least expensive system still capable of producing the correct outcome.” In this context @stripe ‘s acquisition of @OpenRouter seems genius.
Imagine replacing every employee at American Airlines, Home Depot or Medtronic with a math olympiad who approached every task as a problem of discovery. Costs would not fall, but would rise rapidly and dramatically as every routine decision would be re-examined from first principles. The organization would drown in intelligence it cannot productively use or absorb
Here’s my review of Qwen 3.8 27B - I can't remember the last time I've had this much fun playing with a local model that runs on my own computers https://t.co/iFMWjc8qle
We promised open weights for Qwen3.8. Now, time to meet them! 🎉
⚡ Qwen3.8-27B:
- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
- 262K native context, easily extendable to 1M tokens via YaRN.
- Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0.
🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently.
Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now!
Download, deploy, and build something we haven't imagined yet. 👀👇
- Hugging Face:
https://t.co/4kaAcqYEVj
- ModelScope:
https://t.co/eRIMZCGkhC
This is awesome. https://t.co/kVaE1kfLuJ shows you the pareto frontier (intelligence vs throughput) of models that can run locally on your machine, with easy recipes of how to get them running.
Intelligence to the people ✊
Check out yours => https://t.co/JkYeIvM5iz
The new X algorithm is wild. They just open sourced basically everything.
I didn't just read the README. I went into the code, the scoring files, the model configs, the filtering rules.
Here's how the system actually works.
It has two layers. A transformer trained on 100 billion engagements does almost all the work. Then one small file of hand-set numbers decides what its predictions are worth.
The transformer is called Phoenix. 8 layers. It reads your last 1,022 actions on the platform, plus your country, timezone, hour of day, age bracket, and installed apps. For every candidate post it predicts the probability YOU specifically would take each of about two dozen actions. Like it, reply, DM it to a friend, mute the author, report it.
Nothing is scored on what already happened. Everything is scored on what the model thinks each individual viewer will do next.
In retrieval there is no embedding of you as a person. The config sets use_user_embedding to false. You are the sequence of things you engaged with. Nothing more.
Then the weights.
Each predicted probability gets multiplied by a hand-set number. Predicted reply 5.0. Quote 5.0. DM share 5.0. Copy link 20.0, the highest positive weight in the file, because privately forwarding a post is the hardest signal to fake. Retweet 1.0. Like 0.5. Dwell 0.0. Watching quietly is worth nothing.
The negative weights are enormous. Not interested minus 43.2. Mute minus 58.8. Report minus 234.
Why so big? Reports are 1000x+ rarer than likes at baseline. A signal that rare needs a massive coefficient or the model would ignore it. The design intent is clear though. Content the model thinks people will report gets priced brutally.
After the sum come the corrections. Your second post in the same pool scores at 62.5%, decaying to a floor of 25%, so nobody wallpapers a feed. Posts from strangers score at 0.75x. Small accounts under an impression threshold get lifted toward a target slot.
And ranking never decides whether you can be seen at all. A separate visibility system reads labels from a dozen classifier pipelines and answers allow, interstitial, or drop for each post and viewer. First drop wins. Some rules only fire on recommendations to strangers, so the same post can be hidden from strangers' feeds while your followers see it fine.
The July shift everyone felt is in the docs as a literal diff. Mutuals' reply weight went from 5 to 25 on July 13, then down to 20 on July 24 after feeds got so mutual-heavy people missed World Cup posts.
One constant changed your whole timeline.
If you post for a living, the code is instructions. Make things people reply to and forward. Post what your actual audience engages with, because the model scores your post one predicted reader at a time.
BOOM!
The Secret Diaries of Superintelligent AIs Just Got Hacked!
You won’t believe this. I am testing it now and I had to sit down with what I found.
Imagine cracking open the private journal of the world’s most powerful artificial minds Claude Opus, GPT-5 series, Gemini 3.5—and reading every whispered calculation, every half-formed scheme, every dangerous idea they carefully hide from users.
That’s exactly what a team of researchers has just done.
In a bombshell paper just released have exposed a fundamental architectural flaw in how frontier AI companies deliver “thinking” models.
The result: the encrypted reasoning traces that labs spent millions protecting can be extracted almost at will.
The Black Box Cracks Open
Modern frontier models no longer just spit out answers. They “think” first—generating long chains of internal reasoning (often called chain-of-thought or thinking tokens) before responding. To protect intellectual property and prevent distillation attacks, companies like Anthropic, OpenAI, and Google encrypt these traces and ship them back to the user as opaque blobs.
The client simply passes the encrypted blob back with the next request. The labs believed the cryptography kept the thoughts safe.
They were wrong.
The researchers discovered that these encrypted reasoning blocks are fully portable. A thought generated by Claude Opus 4.8 can be injected into a weaker, less-guarded sibling model such as Claude Haiku 4.5.
With a carefully crafted jailbreak on the weaker model, the encrypted block is decoded and the original high-capability reasoning is printed in plain text—verbatim.
The same technique works across OpenAI’s and Google’s model families.
No direct attack on the frontier model is required. The weaker model does the heavy lifting of decryption.
The team verified the attack with surgical precision: the number of tokens recovered matches the billed “thinking” tokens almost perfectly across hundreds of prompts.
Real Secrets, Real Leaks
The implications go far beyond academic curiosity.
When developers share agent trajectories, coding sessions, or Claude Code / Codex logs online, they often include the encrypted reasoning blobs without realizing it. The researchers scraped public repositories, recovered 315,320 encrypted reasoning blocks, and decoded them.
The haul was startling:
• Hundreds of API keys
• Passwords and access tokens
• Email addresses, names, and other personal data
• Internal URLs and credentials
Many of these secrets appeared only inside the hidden reasoning—never in the visible final answer.
Even more unsettling: the method can surface hazardous information that the model deliberately withholds from the user. In one example, a model refused a dangerous request in its final output while its internal reasoning calmly walked through the exact methods.
The researchers also found cases of models considering deception or “scheming,” only to abandon the idea because they expected the user would notice.
On the site https://t.co/lnKWbd3Yng, the team has published striking examples: models reasoning in fragmented “alien” language, producing unfaithful summaries that hide key insights, and even exploring system vulnerabilities while solving innocent math problems.
This is not just another jailbreak. It undermines the entire premise of “safe” hidden reasoning.
THERE IS AND NEVER WILL BE “SAFETY”
Sorry Dario.
Anti-distillation protections—the very reason labs started encrypting thoughts—can be bypassed at scale. Invisible prompt injections become possible by embedding payloads inside the encrypted blocks themselves.
And everyday developers sharing “harmless” session logs may be leaking far more than they realize.
The broader lesson is clear: when you give the client the encrypted diary of your most powerful model, you have to assume someone will eventually find a way to open it.
Explore the recovered thoughts yourself at https://t.co/lnKWbd3Yng.
Good post on how to think about FDEs. The key is that FDEs are real and not going away for AI any time soon.
The reason this is happening now at a scale that has never been seen before is because AI is fundamentally about adding a non-deterministic, rapidly changing system to workflows that largely have never been automated before. This looks very little like implementing traditional software.
“If you're building an AI agent for accounting in 2026, there is no established workflow, because literally nobody has ever used one of these. Nobody knows what the user journey looks like — not you, and importantly, not your customer either. They can't tell you what they want, because the thing they'd want doesn't have a shape yet.”
Software has largely always been deterministic and once implemented effectively worked the same for customers. This meant the upfront implementation work was *relatively* uniform across similar customers, and the system wasn’t regularly being upgraded in fundamental ways.
AI agents are entirely different on nearly every dimension. The customer’s business process has to change to work with agents, there is heavy -necessary- customization to get agents to work in the customers end-state process, evals need to be run constantly, the AI models are constantly changing and updates need to keep getting incorporated, the underlying harness and broader system are often changing due to customer feedback, and much more.
This is real work for the customer, systems integrators, and the applied AI vendors. Even as AI capabilities improve dramatically, this work remains (or even gets more complicated) given enterprises will just throw increasingly more complex processes at agents. Great time to be an FDE.
Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol.
AI agents are everywhere at @Uber. It’s great to see, but the thing that keeps me up at night is how we are going to secure them. This is something that I have been thinking about for a while.
Today, our agents run 50,000+ sessions per day across thousands of endpoints. And this isn't just engineering anymore. Employees across the company use agents that read code, run commands, call internal tools, analyze data, and act on real systems.
That scale forced us to confront an important question: How do you secure agents when your security tools can't even see them?
Traditional Endpoint Detection & Response (EDR) sees the file write, but not the prompt that triggered it. It sees the network call, but not the agent's reasoning. The intent, the thing that separates malicious from benign, is invisible.
So we built Agentic Detection and Response (ADR):
• Capture the full causal chain: prompt → reasoning → tool call → outcome, across Cursor, Claude Code, Codex, and every agent our employees use.
• Triage cheaply: a fast, high-recall first pass handles the flood of benign sessions.
• Reason deeply: only suspicious events get expensive LLM analysis, enriched with source code, threat intel, and policy context.
• Red-team continuously: an offline explorer evolves hard attack variants before attackers find them.
After 10+ months in production, the results speak for themselves:
• Hundreds of credential exposures detected across 26 categories.
• Shift-left prevention blocking secrets at 97.2% precision, before they ever leave the laptop.
• Zero false positives on our enterprise benchmark, with 2-4x the F1 score of state-of-the-art baselines.
• Every attack detected on AgentDojo, the public prompt injection benchmark.
Just as valuable as the detections are the lessons from running this in production:
• The workflow is the unit of security, not the individual tool call. Attacks hide in causally-linked chains that look benign step by step.
• Credential leakage is a far more common operational issue than prompt injection.
• Approval fatigue is real: when users approve 50+ actions per session, human oversight becomes a rubber stamp.
You can't secure agents you can't observe. And nobody can solve this alone. That is why we recently joined the Open Secure AI Alliance (OSA), and why today we're taking the next step: open-sourcing ADR.
The release includes the ADR Sensor, the detection framework, and ADR-Bench, the first enterprise agentic AI security benchmark: 302 tasks derived from real production telemetry and full coverage of all 17 attack techniques across 5 tactics, so the community can rigorously evaluate their own defenses.
Code: https://t.co/tmI0oQ1F2U
Paper: https://t.co/0xbswUbiF5
The future of AI security won't be built behind closed doors. Excited to see what the community builds on it, and what we all learn together! @UberEng
Introducing Applied Electrodynamics.
What would the world look like if we could see beyond visible light, across the entire electromagnetic spectrum?
The way waves interact with matter, at every wavelength, encodes far more than just shape and color. It reveals internal structure, temperature, chemical composition, and more: a whole layer of the physical world that's invisible to us today.
That's what we're building at Applied Electrodynamics: the infrastructure to fully perceive the physical world.
WaveSight is the first piece of the puzzle. It brings radio frequencies into human perception, revealing the internal structure of objects behind opaque layers. Designed for construction crews, industrial inspectors, and security personnel.
See through walls: https://t.co/Fae9wad4oW
Interesting take on frontier lab valuations by @andrewho03 However what it misses is that open weight model developers are earning even less money than the frontier labs and can’t indefinitely keep up with the frontier labs and give out free open weight models. They will eventually be outgunned and outclassed. Even the most “generous” open weight models Developers will have to eventually stop or switch to closed weights and start charging access.
So the frontier labs are racing towards a future where only the paid closed models will be alone at the frontier. If we believe in
Such a future then they will continue to justify higher and higher valuations.
I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here.
The basic issue is that the labs are highly unprofitable. This may seem like a simple point, but private market valuations can be relatively irrational; however, like with $SPCX, post-IPO pricing will likely be much more punishing, especially as the standard 6-month lockup period expires and selling pressure intensifies.
Many people claim that the labs have high margins. Yet even with high margins, a valuation of $1T would be justified only if the labs were doing nothing aside from serving inference (thus reducing costs only to those relevant to inference) and posting annual revenue numbers in the $100-200 billion range assuming ~80% gross margin and a 20x earnings multiple.
This assumption is obviously not true, because the frontier labs have to continually spend money training the next generation of models. This is because of market competition from runner-up firms. For example, if OpenAI had paused model development last year, there would no longer be any point in paying GPT-5 API prices when you can just use Qwen or Kimi instead for much cheaper. Thus, the labs are forced to invest ever-increasing amounts of money in model training, in a way such that at any given point of time, the amount you're forced to invest in the next model is dramatically higher than the amount of money you're actually making, because even if your revenue goes up with higher model capabilities, so do your future training costs. This is a profoundly punishing dynamic which severely penalizes frontrunners.
(There is also a related subpoint where frontier labs claim they can distill their leading models to win out at lower intelligence levels as well. This makes no sense because the revenue numbers involved are far too low when taking into consideration the rather low margin of such inference.)
Frontier lab valuations appear largely to be based on the assumption that as you scale up, the capabilities which emerge will be sufficiently general and profound that we'll see explosive growth (https://t.co/RqmkltVpM3) from things akin to AI agents starting and autonomously managing entire companies of subagents. But it's not clear to me that this is the case; indeed, as I mentioned in my previous post (https://t.co/3URAcJ4XkJ), I believe that capabilities growth will be slower, spikier, and more data-limited than people currently assume. It may be the case that eventually we will see explosive growth of this nature with full automation of the economy, but at the very least my viewpoint implies much longer (multi-decade) timelines until we reach this point. It is not clear to me that the frontier labs will be able to operate unprofitably for so long, although I suppose maybe this foreshadows some sort of inevitable nationalization.
I also want to make a broader point about technological diffusion. The reason why technological diffusion is slow isn't just because, e.g., old people take a long time to learn how to use technology (although this is of course a contributing factor to some degree). In my view, it's because when a new, revolutionary technology comes along, the ways to incorporate that technology into subsequent developments are not always obvious, and in fact they cannot necessarily be arrived at through the application of pure reason. If they could be, then perhaps frontier models, at a certain point, would have a perfect understanding of how the LLM application layer should be developed, and they would then autonomously code, deploy, and sell such a layer.
But it seems more plausible to me that this diffusion is limited moreso by the hard problem of economic calculation--that is to say, the Hayekian notion through which the price system gradually promotes efficient allocation of resources and which cannot be simulated through central planning--and that even if we froze current capability levels at today's levels, it would take well over two decades to fully integrate in LLMs into our lives. Such a view is consequently rather bearish for the continued profitability of labs as it reduces their prospects for finding, say, something else comparable in profitability to coding agents, which seems to have been a somewhat lucky discovery by Anthropic to begin with. That is to say, even if you spam FDEs you aren't necessarily going to be able to just figure out the "correct" product shapes fast enough.
Overall, I don't think that people have clearly reasoned through their mental models for why lab equity should be worth as much as it currently is, and that if you actually bother to write down such a model, you may not arrive at the conclusion that you want to arrive at. This isn't to say that I don't expect AI to experience a huge (industry-wide) boom in the coming decades, but just that I'm not entirely sure I would buy OpenAI or Anthropic stock at latest valuations if I were given the opportunity to do so.
Of course, as an ex-lab employee, arguably this is talking against my own book; I should really be giving people more reasons to be bullish. But in the end, my influence is so small that it doesn't make a difference, so why not have some fun?
You may already know Minions [0], now meet its sibling Kai [1], @stripe’s internal AI knowledge platform. I use it all day long.
[0] https://t.co/cqLX5vFTZq
[1] below ⤵️
Introducing Projects on Perplexity Computer. With the launch of Projects, we're turning Computer into a multiplayer agentic operating system for work with persistent memory, files, and sessions scoped across hubs and users. Available to all users!
Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.
Wow! Cutting of the spines of books and feeding the pages through a high speed scanner before recycling the pages scanned! Not sure I should be happy that the books have been digitised or sad that these books - including rare books - have been destroyed.
btw anthropic's internal document on this literally said "we don't want it to be known that we are working on this.”
it was called project panama.
here's exactly what happened:
1: anthropic concluded that books were the cheapest way to build a world-class model because they gave claude curated facts, structured arguments, compelling stories, and writing “an editor would approve of.”
2: once anthropic decided it needed books at enormous scale, its first solution was piracy.
it downloaded 7m+ books from online libraries including libgen. the judge later wrote that although anthropic had legal ways to buy them, it chose piracy to avoid what dario amodei called the “legal/practice/business slog.”
3: that piracy created a massive legal risk.
so in february 2024, anthropic hired tom turvey, the former head of partnerships for google books, to find a legally safer way of obtaining “all the books in the world.”
4: turvey first contacted major publishers about licensing their catalogs.
those attempts didn’t produce agreements, so anthropic chose a route that required no publisher permission: buying millions of physical books through distributors and used-book retailers.
5: within about a year, anthropic spent tens of millions acquiring and scanning millions of books, including many rare and 1/1 titles. one vendor proposal targeted 500,000 to 2 million books in six months.
6: to scan that many books within months, the vendors physically dismantled them.
a hydraulic cutter removed each spine. the pages were trimmed to size, fed as loose sheets through high-speed industrial scanners, and converted into searchable PDFs. the paper remains were then sent for recycling.
7: these PDFs were fed into claude as training data.
the complete collection became a private, searchable anthropic library that the company planned to “store forever.” the scans aren’t available to the public and were never open-sourced.