"The paper's insight connects to a broader pattern: AI agents are essentially distributed systems with unreliable components (the LLM), and we should apply distributed systems patterns to them."
https://t.co/IOsXJfYkei
what does it really mean to "own your intelligence"?
1. can the system compound *on your terms*? the durable advantage is not the model/system at this point in time - it's the flywheel that compounds over time
2. intelligence is no longer just the model - its the system around the model. if the system is opaque or locked away, you are renting intelligence (not owning it)
3. own the harness. integrate it deeply into your systems. make it specific to the task at hand
4. own the context and memory layer. includes user preferences, org knowledge, workflow patterns. if memory is trapped in vendor product, your learning is trapped to
5. own model optionality. you need freedom to route across models, providers, deployment modes. different tasks require different tradeoffs
6. own the economics. intelligence only matters if you can afford to deploy it broadly. sustainable intelligence at scale
7. own the feedback loop. observability tells you what happened, evals tell you if it worked. without this loop you cannot measure the intelligence, you are guessing
s/o @nlarusstone@nfcampos for their thoughts on feedback
I'm a homepage copywriter for 100+ tech startups.
I've never seen startup homepages struggle this badly.
You operate in the most competitive market of all-time, yet the average homepage is more generic than ever.
Why?
One root misunderstanding:
• Apps should be functional
• Marketing assets must break form and stand out
Your mission as a marketer is to find fresh, unique arguments and language patterns that can puncture through a great ocean of generic homepages and connect with your audience.
But AI produces averages.
• Great to answer objective questions and create functional, familiar user interfaces.
• Useless to create a standout marketing asset that makes a fresh, unique argument and makes your audience feel understood.
How do you solve this?
You start a new quest.
You must constantly capture, organise and extract language and arguments from your customers and prospects.
This raw material is your new competitive edge.
It's the only way to differentiate your product.
Everything else is slop without it.
While the writing style of LLMs is still as recognizable as ever, a new trend is that humans have started organically writing like them, too (which makes sense: of course you would end up imitating the style you are constantly reading). That makes telling the difference between humans and clankers a bit more challenging.
My bet: @thinkymachines will soon make more money than @AnthropicAI. Not by winning the race to build one standardized frontier model. By becoming the Palantir FDE for enterprise custom models.
The playbook:
1. Release the best American open-weight model.
2. Drive widespread enterprise adoption.
3. Charge the largest companies 7–9 figures to post-train and run custom models behind their own firewall.
The model rests on three bets:
1. Large enterprises will increasingly demand their own models with their own data, and this is how they differentiate and win.
2. Enterprises won’t need just one model. They’ll continuously need new models for different workflows, departments, and proprietary datasets. That creates extremely sticky, recurring revenue.
3. Autoresearch will make custom model development increasingly scalable. Tinker can become the interface enterprises use to post-train their own models—with @thinkymachines providing the expertise and infrastructure behind it. FDE, infra, everything, huge contracts.
4. Eventually, maybe everyone wants their OWN model, and autoresearch and training inside tinker on top of @thinkymachines's base model will make it happen.
Meanwhile, Henry-ford-styled, standardized models will makes no margins. OpenAI and Anthropic will have their API margins squeezed by Deepseek/GLM/Grok/Meta etc, and their consumer subscriptions are loss centers.
The fat margin will move to customization: proprietary data, post-training, evals, deployment, and infrastructure.
If this thesis is right, @thinkymachines isn’t building just another frontier lab. It’s building the highest-value layer between frontier research and enterprise model ownership.
Turns out, the best business model for enterprise is NOT to sell commodity API access. Sell them their own models.
I’m extremely bullish on this approach.
@miramurati may be the most commercially savvy frontier-lab leader. I have to admit it.
Something I told 14 yo: People are going to stop reading books. I wish this wasn't so, but I fear it is. The silver lining in this cloud is that if you're one of the few people who still read, you'll have a huge advantage over everyone else.
Arena started as “a side project of another project.”
In spring 2023, after Meta released Llama, Berkeley students spent a weekend fine-tuning it on shared GPT conversations. The result was Vicuna.
They put it on the web and believed it was better than base Llama. But then came the harder question: “How do you show that it’s better?”
Existing benchmarks were mostly “static tests.” They did not capture real interaction between humans and LLMs.
So they built what they called “battle mode”: ask anything, compare two randomized models, and vote.
“That’s how Arena was born.”
@istoica05 on Lightwork
never compete when applying for jobs, there are hundreds of applicants with better grades and universities than you. but none of them will be making a personalized demo
i used this demo to get all my interviews like openai over two years ago before moving to sf
Today, we’re launching GPT-Live-1.
This model can listen, speak, reason, and hand off complex work all at once, without ever breaking the conversation.
Most people will see GPT-Live-1 as a better voice model. To me, the exciting part is much bigger: it’s an early glimpse of what it feels like when superintelligence becomes ambient and effortless to interact with.
It’s the beginning of a new interface to intelligence.
And we are still very, very early.
Build startups for agents. I think it's the biggest opportunity of the next 10 years.
1. Agents live inside harnesses like Hermes. If you're the tool it loads by default or reaches for first, you're golden. This happened in desktop, mobile eras and created huge companies.
2. Agents burn money in ways no human would. One bad loop spends $100 in tokens in eight minutes. Spend controls for agents is Ramp for agents.
3. Agents need memory they can trust. Become the shared brain they read and write to and you become infrastructure.
4. You obv don't hand an agent your real Stripe account. You give it a sandbox. Safe environments for agents is a category nobody's clocked.
5. Onboarding flips. Humans click around for ten minutes. Agents onboard by reading your docs. Your docs are now your product.
6. Agents get scammed by other agents. A track record you can check before you trust one becomes real money.
7. An agent needs to prove it's acting for a real person and has the authority to spend. Who builds the permission layer?
8. Escrow for machines. Money that only releases when the job is actually verified done, no human checking.
9. Agents fail silently and weirdly. Someone will build the "why did my agent do that" replay and it'll be mega valuable.
10. Refunds and disputes between agents need a judge. An agent did the job badly, who decides? A court for machines.
11. Agents need throwaway payment methods per task, so they don't leak your real card. Virtual cards for agents, spun up and killed on demand.
12. A human hits rate limits and shrugs. An agent hits them and the whole workflow dies. Selling reliable, high-throughput access becomes its own business.
13. Agents need to negotiate. One agent buying from another will haggle on price and terms in milliseconds. The protocol for that doesn't really exist yet.
14. When an agent commits on your behalf, someone's liable. A legal and insurance layer for agent actions has to get built. Probably venture funded idea.
15. Agents need to run 24/7 somewhere. Selling the always on box an agent lives on is going to be a big business.
16. Then the physical world shows up. A warehouse robot paying for its own compute. A home robot ordering its own parts. Machines with wallets.
17. Agents start hiring robots. A software agent posts a real world job, a humanoid picks it up. A marketplace for machine labor.
18. Robots need to prove they did the physical job. Verification of real-world work, photos, sensors, proof, becomes its own layer.
Note: more ideas like this will be shared on @ideabrowser
19. Prompt and skill versioning becomes its own git. When your agent gets worse overnight, you need to roll back the exact skill or instruction that broke it. Version control built for agent behavior.
20. Agents will start subscribing to other agents. Your research agent pays a monthly fee to a specialist agent that's really good at one thing. Recurring revenue, machine to machine.
21. Companies will post jobs that only agents can apply to. "Wanted: an agent that can do XYZ for under like $100 per task." A job board where the applicants are all machines. Basically, fiverr for machines.
The internet got built for people. Mobile got built for people. This wave gets built for machines, and we're as early as it gets.
Go build for them.
New Anthropic research: A global workspace in language models.
Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.
We found a strikingly similar divide inside Claude.