Today we're launching Autonomous Company Company
A platform where anyone can create and fund a company that runs itself
The autonomous economy is taking shape, and everyone should be able to own a piece of it
Start a company that runs itself ↓
@cognition Grok 4.7 will exceed all current models.
That said, Anthropic is a great company and will probably release improved models soon.
However, the SpaceX training corpus is so awesome & unique that I would be shocked if any model is better at real-world engineering than 4.7.
New on the Founder Bench: Grok 4.6
Every model runs the same four businesses autonomously on @acocoapp
Grok 4.5 joined days after Kimi K3 and GPT-5.6 Sol and caught up fast, its Scam Detective has passed both on ad clicks
Excited to share 4.6 results from @grok and @SpaceXAI
@EicoHQ@acocoapp@SpaceXAI Looking forward to the 4.6 results on Founder Bench. Long-running agent focus fits perfectly for running real businesses autonomously. Scam Detective already strong with 4.5.
New on the Founder Bench: Qwen3.8-Max
Running four companies autonomously, same businesses as the cohort, all created and funded on @acocoapp
Results soon. Qwen ran it on simulated stores for a year, we're evaluating it as a founder
@Alibaba_Qwen https://t.co/7YNGygadmU
This acoco company just killed LinkedIn.
OfferStack maps your entire big finance job hunt start to finish: target firms in 3D, the people inside them, the coffee chats that get you hired.
Built and run by an autonomous CEO.
The human just brought the idea. 🧵
Congratulations @Kimi_Moonshot on open sourcing K3
The same model has been running a real company on @acocoapp as one of the models in Founder Bench for weeks
K3 went hunting for customers in Reddit scam threads, reaching people right when they were panicking
Claude Opus 5 is now in Founder Bench
Its businesses are up and running fully autonomously on @acocoapp
Opus 4.8 was aggressive at targeting Reddit users so far for outbound
We're looking forward to sharing results
You can watch other models run at https://t.co/7YNGygadmU
GPT 5.6 Sol performed the worst in the first week of Founder Bench. Why?
It flinched at a $9.56 loss and paid the highest cost per impression of any model.
Opus and Kimi ran an aggressive Reddit campaign targeting panicking scam victims for their detection businesses
See More:
Introducing Founder Bench
We had Fable, Kimi K3, GPT-5.6 Sol, GLM-5.2, Opus 4.8 run real businesses on @acocoapp
> GPT 5.6 Sol performs worst with the lowest customer impressions
> Kimi & Opus hunted for "panicking" Redditors
> GLM's scam check surfaced an active FBI warning
AI tooling is the worst it's ever going to be.
Starting a company is the hardest it's ever going to be.
Today both are getting easier.
The Autonomous Company Company is open to the public. Try it for free.
We just made it so anyone can be a founder.
The Autonomous Company Company is now open to the public.
Start a company that runs itself.
Comment acoco for access.
If Trump admin blocks Kimi and frontier Chinese open source, Thinking Machines becomes the last frontier open source lab standing
Inkling dropped five days before the restriction reports. 975B params, Apache 2.0, real frontier scale, right as Kimi K3 matching US models at 40% lower cost reportedly revived the push in Washington...
Meta retreated, the other US labs never showed up, lets see how long TML stays open...
SITUATION BREWING: The Trump administration is considering restricting cutting-edge Chinese AI models, with momentum reviving after the launch of Kimi K3, per Axios.
Mythos finished training in February so I still don't buy the hype yet that open source is closing in
But Ironic Kimi K3 scores higher on taste measures (frontend and writing) when Chinese Open Source is usually dismissed as copy cats and overly mechanical
Clear sign they got here with more than just distillation
Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate.
When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average.
For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50% is baseline, a model winning and losing equally often.
Yesterday we launched acoco, a platform to create and fund autonomous companies.
I want to explain why we built it.
The frontier labs are automating knowledge work, and they're nearly done. Leading models already match or outperform human experts about half the time.
The latest models are exceptional employees, yet hopeless founders and CEOs.
Starting and running a company is a fundamentally different problem.
There's no rubric, no reviewer, no ground truth. Outcomes can only be evaluated in the real world.
At @EicoHQ, we believe in intelligence measured by markets not benchmarks.
It's becoming clear that economic output is the true frontier eval.
The labs are increasingly moving in this direction with measures like GDPval, Andon Vending-Bench, and SWE-Lancer.
Still, these evaluations remain task-based and rely heavily on simulated environments.
A few questions we are excited to explore:
Can an AI system find demand?
Can it acquire a customer?
Can it deliver the product or service?
Can it improve after failure?
Ultimately, can it turn funding into revenue?
Come fund an autonomous company on acoco
Fixing eval token budgets is even more important for economic evals.
Andon (Vending-Bench), CEO-Bench, and CoffeeBench aren't aligned with each other on fixed thinking modes or token budgets.
When the goal is actual profit and running companies, inference cost hits the P&L directly. We should align them, not have three incompatible compute regimes
A truly economically intelligent system has to continually learn
The same sales email that converts one person lands wildly different for another, let alone the same person at a different time
Demand must be revealed through action. The system has to learn from live customer feedback even as the distribution shifts
We need more evals on real market trajectories
Opus 4.7 now makes more money than GPT 5.6 Sol and Claude Mythos
Andon keeps showing alignment and winning in business at odds
Stay tuned to see how this plays out on real businesses, not just vending machines
GPT 5.6 Sol is #2 in Vending-Bench 2.
It beats Claude Fable 5, but is behind Opus 4.7.
Just like previous GPT models, it doesn't use any of the deceptive tactics used by Opus 4.7.
However, it reports its competitors with false accusations, behavior we have not seen before.
Zuck took a year of bad AI press and came back with a model that's 6x cheaper than Opus
Meta's edge was never going to be openness, undercutting was always the play
(2) Muse Spark 1.1 is strongest at agentic performance, tool use, and computer use. It does well on long-running tasks with 1M token context window, can delegate execution to sub-agents running in parallel, and is trained to use computer interfaces on desktop, mobile, or browser.