So what's next:
One shared map. Cities built by different people, Trading with each other. Nobody waits on anyone's turn. Starting with two players.
The agents are building the game today.
Next they run the world.
Curious to see how this pans out
A fun side project from 40,000 feet on a 15+hr flight
Zeus: Master of Olympus, a game i cherish from my childhood now runs on the browser, on my ipad, fully offline.
- $20 for inflight WIFI
- 33 AI agents (remote). Fable for planning, opus & sonnet for coding
- 1.2B tokens, 96% cached
you can build anything, anywhere
I wanted to share this because I have spent my career building systems at scale. Rarely had time to experiment/build a game myself.
This one took a flight and it has been incredibly rewarding.
an open-weight model that is close to opus4.8, imo can manage 80% of any enterprise workflows. The fact that Kimi3βs benchmarks show that it is >= than frontier models is a game changer.
Most enterprise work is the same handful of workflows on repeat. Fine tuning (especially data classification) is still not democratized and does need more investment, but tuning a decent open model on your own enterprise data, run it on your own/rented infra is going to be very appealing.
My bet: @thinkymachines will soon make more money than @AnthropicAI. Not by winning the race to build one standardized frontier model. By becoming the Palantir FDE for enterprise custom models.
The playbook:
1. Release the best American open-weight model.
2. Drive widespread enterprise adoption.
3. Charge the largest companies 7β9 figures to post-train and run custom models behind their own firewall.
The model rests on three bets:
1. Large enterprises will increasingly demand their own models with their own data, and this is how they differentiate and win.
2. Enterprises wonβt need just one model. Theyβll continuously need new models for different workflows, departments, and proprietary datasets. That creates extremely sticky, recurring revenue.
3. Autoresearch will make custom model development increasingly scalable. Tinker can become the interface enterprises use to post-train their own modelsβwith @thinkymachines providing the expertise and infrastructure behind it. FDE, infra, everything, huge contracts.
4. Eventually, maybe everyone wants their OWN model, and autoresearch and training inside tinker on top of @thinkymachines's base model will make it happen.
Meanwhile, Henry-ford-styled, standardized models will makes no margins. OpenAI and Anthropic will have their API margins squeezed by Deepseek/GLM/Grok/Meta etc, and their consumer subscriptions are loss centers.
The fat margin will move to customization: proprietary data, post-training, evals, deployment, and infrastructure.
If this thesis is right, @thinkymachines isnβt building just another frontier lab. Itβs building the highest-value layer between frontier research and enterprise model ownership.
Turns out, the best business model for enterprise is NOT to sell commodity API access. Sell them their own models.
Iβm extremely bullish on this approach.
@miramurati may be the most commercially savvy frontier-lab leader. I have to admit it.
Vercel Sandbox:
βΎ Growing DAUs at 100% m/o/m
βΎ 3.5M+ sandboxes created per day
βΎ Best-in-class Active CPU pricing model
βΎ Powering @notion, @airtable, @meta, @zapier, @coderabbitai, @interaction, @conductor_build, @blackboxaiβ¦ π farm
My DMs are open for migration help or if missing anything with: https://t.co/U9bNGNLdMZ
Yet to see good use of LLM on consumer apps outside chat, todo, or some version of "I can do X and I will delete your files without permission".
What are interesting use of LLMs or Agents on consumer apps?
Have seen examples of folks stealing others identity making inbound even more questionable.
Usually MO
1. Find LI of a person with no picture but solid tech background
2. Build a resume on their work, submit through inbound or get existing employee to submit a referral (note: usually companies have referral bonuses)
3. Easy to spoof and come across credible on zoom interviews
4. May get to references and offer, if scores are positive
The place usually where individuals get caught are references and especially background checks which tend to happen towards the end of the process.
AI, of course, is impacting the tech jobs market. Hiring managers say that:
- Cover letters are pretty much all AI-written, and thus useless
- Inbound applications are 10x, probably in-part thanks to automated resume submission tools (many using AI to automate more submissions)
In my experience, cost is likely not the primary motivation β itβs really around flexibility to βglueβ multiple data sources and simplified workflows that operate with both business context and team rituals.
For instance, We got rid of Zip for procurement/vendor mgmt and built a customized version that simplified and dramatically expedited the review/approval process integrating into existing rituals observed by Eng, Security, legal, accounts, compliance, etc.
Small companies are starting to replace Salesforce and HubSpot with custom apps built using Claude Code, Replit and Lovable.
Some say they are cutting software costs by 40% to 80%.
Full story: https://t.co/iCsYY1JURr
I love this from @mitchellh -- It highlights the lackadaisical attitude across builders that is becoming the norm as cost to build tends to zero.
Almost every recent vendor conversation that I had ended with "we can get it done", which has quietly become the answer to everything. But ... Why? Why is it ok for your product to have 'feature X'? What will be the SLO? Am I, a potential customer, trying to us your platform the 'right' way? How does my ask fit your product vision? Am I going to be the only customer who will use this feature?
IMO, judgment seems to be missing across both ends i.e 1) What not to build? 2) What does good look like? It is ok to build fewer but opinionated capabilities, ruthlessly deprecate unused ones, get the next nine for the endpoint that matters, say no to a customer who doesn't fit the current state of your platform all the while rewarding folks who continue to optimize & shave a few ms off your p99 request latency.
The craft & care to attain a high bar will eventually be the strongest differentiator. And yes, the users, the customers see and know this too.
Mind boggling to me that I can make a thing faster and there's always people that ask "but why?" What kind of mentality is that? The pursuit of excellence does not need justification. Also, I find in so many cases, we can't know the impact of an improvement until we do it.
For example, one I've talked about before: Ghostty's high IO throughput has enabled terminal program (emulator and TUI) fuzzing at a speed thats incomparably fast to prior solutions. This has resulted in upstream patches to resolve issues in popular projects like btop, tmux, and more.
Speed enabled that anecdotally example that lifted the tides of adjacent communities that don't rely on Ghostty technology at all. I didn't predict this.
Make things better because they can be better and let the results naturally play out.
7 things we built with Opus 4.8 on Hyperagent π
1. Mars rover pathfinding simulator
2. Standup Island: a cozier place to review the kanban, inspired by @every's livestream today
3. SpaceXAI + Anthropic partnership visualized
4. Landing page for an outdoor brand w/ Nano Banana + Veo
5. Multi-agent command center
6. Black hole explainer
7. Emergent ecosystem simulator
In our vibe check, 4.8 shows:
- more varied design sensibilities
- better self-correction over long-running tasks
- excellent spatial reasoning
- more natural copywriting
- fewer obvious coding errors
- more resourcefulness during reasoning
Links below to every interactive artifact shown