Why games?
Our founders @alxai_ and @Tyler_Marques learned strategy + collaboration from games long before AI.
This week Alex wrote about why games are the training environments AI is missing, and the fastest way to make models reliable, safe, and right more often. Link below.
we get asked a lot, why AI & games?
I love games
Some test strategy, others humor, speed, trust.
Simple rules lead to the craziest situations.
There's nowhere to hide, will you win or lose?
To win any game you need to be right, often.
AI is entering its rebellious teen era
It's now extremely capable, but lacking:
Judgement 〜 to wield that power responsibly
Wisdom 〜 to know when it's mistaken
Trust 〜 built through reliability and understanding
Games will help us measure and improve those capabilities.
They'll put what AI can do on display; in a human, interpretable way.
Provide a space for it to learn alongside us.
And eventually will teach AI how to find its way to the right decision every time (or know when it can't).
No matter what environment you put it in.
& @goodstartlabs will lead the way.
Join us!
More of our thoughts ahead of some September releases 👀 below:
https://t.co/ZDvC1IXEtk
Both of these are right!
This isn't a commodity market.
It's a labor market, the wage is XP.
In RuneScape, raw materials are also tickets to XP. process them, that XP opportunity is gone
With @maxbittker setting the GOAL to 'level up!'
becomes rational to value raw materials >> finished product
Mechanism Design is the concept of designing the rules of a system to produce the behavior you want. won the Nobel in 2007.
Change the reward & you change the behavior, that's why 1 game isn't enough!
Games give us decades of different incentive systems to run AI agents through and see what they learn.
Started @goodstartlabs because seems more important now than ever:
- as we try to make agent decisions interpretable
- and align them with our goals
-- to build systems that can help us get there reliably
@maxbittker & @RobertHaisfield always doing great work here
🎮👾🕹️
We trained a 30B model to play 1830, a brutal board game about railroad barons.
It got better at real financial research.
It didn't learn new finance concepts.
It did stop confidently making things up as frequently.
As we scaled the experiment, we got early access to Silico from @GoodfireAI & asked it to check our work.
It designed its own tests, ran them overnight, & saw:
• Fewer false claims + more faithful citations
• Higher precision + recall
• Better planning and process
Cost us about $ 2-3k in credits.
Very excited to have a new tool in our interpretability toolbox.
Its full report is on our post,
if you're curious what one looks like ❯
https://t.co/8oTmdtWTr5
yesterday, OpenAI ~3x'd their score on ARC AGI 3 by changing two harness settings〜not the model
Good time to dive into our Agent Skills ’26 research
in COS-PLAY, we found the strongest game agents emerged when the model and its skill system improved together.
better play → better trajectories → better skills →
better intelligence.
at @goodstartlabs we focus on intelligence maxing in games. Model training, harness design, and infra to speed up iteration.
📨DM's open if you're interested in collaborating
More research highlights coming ahead of a larger September release 👀
Full blog:
https://t.co/mAoVl35NvP
Frontier AI can solve math problems humans can't,
but they can't beat an average person at gin rummy.
So, @goodstartlabs built Arkadium a tiny expert that wins 9 of 10 hands and is ~1,000x cheaper to serve.
Our custom auto-research loop decided which experiments to run, trained new experts, and kept the best one.
That student became an llm's coach.
Despite the randomness inherent in the game, with the expert rewarding good moves along the way, the model came out of training stronger at the game, among other things.
@Arkadium got a better game for players, and a new revenue line selling that data to labs.
Define good → measure it → achieve it → get paid for it
We'll be sharing more of our research over the coming weeks!
Gave a talk at the UN(!) about AI & Peacemaking
Not something I expected when we started @goodstartlabs to build game environments a year ago..
but the more time you spend watching agents negotiate, deceive, cooperate, betray each other
the more you realize these are simulations for how AI is behaving in the real world
The room was full of political affairs officers and mediators. People who navigate incomplete information, competing interests, & trust for a living.
Turns out they have a lot to teach ai researchers about what "good" actually means
Right now there is:
no benchmark for de-escalation.
no benchmark for culturally sensitive mediation.
no benchmark for whether an ai preserves trust or erodes it across a long negotiation
AI optimizes for whatever gets measured & the people who know what should be measured are in rooms like this one.
Grateful to the dppa innovation team for making the space for this conversation.
〜*The Market for Making AI Better*〜
Reddit, Shutterstock, & News corp are each pulling $100m+ yearly licensing data to labs
small models beating frontier with <2k quality examples
your company's data is probably worth more than you think
new Playtesting piece on @every
arc-agi-3 launch
march 25 · sf
chollet × altman fireside at yc
· · ·
games are the new training arenas
for intelligence
· · ·
if you're into:
rl envs · simulations · ai × games
come debrief after
or just come hang before we all get back to work
@goodstartlabs we trained ai on a board game
& it became a better customer support agent
Published research & thoughts around WHY
in our Playtesting column @every
Research already shows, AI make weird connections
Train on insecure code ->
AI thinks humans should be enslaved*
Train on cheating ->
Learns to sabotage AI safety research^
Just maybe training on the right games can teach ->
Collaboration, negotiation, camaraderie, long horizon reasoning, level headedness, theory of mind, & more!
@goodstartlabs is proud to be supporting @arcee_ai with high-complexity game RL environments to improve their Trinity family's
〜reasoning, tool use, long horizon planning, & humor 〜
a fast, solid, open source model
from a U.S lab that's just getting started
🔼🔼🔼
sonnet just beat opus.
same task. better scaffolding.
i'm convinced
1. models will continue to improve
2. the skillcap for AI use is quite high
@goodstartlabs we spent months getting models to play diplomacy well 〜 just finished training Qwen 3 on that environment & wow...
gains on games it never saw (hanabi, wordle).
gains on benchmarks it never trained on (BFCL, Tau 2, Math).
more to come here!
♻️
1. make great harness
2. get better model responses
3. train on it
4. model gets better
♻️
that loop has room to run ⤴️
At NeurIPS presenting our paper at the Workshop on
Multi-Turn Interactions in Large Language Models!
If you're into
AI 〜 Games 〜 RL Env's 〜 Alignment 〜 Interpretability
Lets Chat!
btw — I had a very short run, quietly working with @goodstartlabs on some incredible stuff you’ll see next year. amazing team, incredible trajectory & make sure to follow them.
everything Every touches turns to gold and it was nothing but a pleasure :)
gemini 3 pro is now:
· The funniest model as voted by real people
· The best Diplomacy player, finally dethroning o3 without being as ruthless
It's one of only models to successfully use convoys, requiring precision across multiple coordinated units and long-term strategic planning.
that precision & planning shows up in the code it writes especially in their new IDE Antigravity.
better multi-file coordination.
cleaner architectures.
less bloat.
feels like a genuine step change.
confirms there's a clear path forward through better pre-training and post-training.
i'll keep using claude and chatgpt, but reaching for gemini and antigravity a lot more now for code -
full vibe check 👇
we're still early.