👾Working in AI, connecting dots between people, insights, & odd little ideas.
🎧Love music, learning to make it, too many side quests.
🫧A restless being.
Interesting work. Persona fidelity may matter far beyond simulated-user evaluation.
The research reminds me of a question I’ve had about multi-agent systems:
when agent collaboration breaks down, could part of the problem be that the agents are not actually diverse enough?(I vaguely remember this coming up over dinner with @istdrc …?
Human organizations can work in very different ways, partly because the individuals within them are different.
For example, Rick Rubin once described three types of bands that can all produce great work:
- Tension-driven, like the Beatles. The contrast between Lennon and McCartney created friction, but also creative breakthroughs.
- Leader-led, like Tom Petty and the Heartbreakers. Everyone was highly capable, but one person ultimately made the call.
- Consensus-driven, like U2. Everyone had to agree before the band moved forward.
None of these structures is inherently better. Each works because it fits the personalities involved, and because those personalities learn how to work with one another.
Today, we can give agents different roles and instructions, but they may still lack stable, genuinely different personalities. Without that diversity, it may be difficult for them to develop distinct and reliable ways of collaborating.
That’s why the 91.5% persona-adherence result is especially interesting to me: if models can reliably maintain different personas over long interactions, this could be useful not only for simulating users, but also for building better agent teams.
Harvard and MIT just built an AI simulation containing 8.3 billion virtual people, roughly the entire population of Earth.
They call it MatrAIx.
It is a population-scale digital simulation infrastructure powered by frontier models like GPT and Claude.
They mapped out an insanely detailed dataset called Persona 8B. It contains 8.3 billion unique digital profiles.
Every single persona is defined across 1,290 categorical dimensions.
Backgrounds, psychology, spending habits, behavioral quirks, technical literacy, and lifestyle.
They didn't just build a database of stats. They brought them to life.
The researchers dropped these billion persona agents into four distinct digital environments:
• Surveys
• AI chat interfaces
• Live web browsing
• Native desktop and mobile apps
Then they tested how this digital Earth reacted to real-world products, software updates, and pricing shifts.
The system recorded everything: how long an agent hesitated after a price increase, when they abandoned a broken checkout flow, and how much latency they would tolerate before closing an app.
When researchers tested persona adherence, the AI agents successfully maintained their assigned psychological and demographic behavior 91.5% of the time.
Think about what this means for the future of business.
You no longer have to wait weeks for a market research firm to tell you if a product feature will fail.
You can simulate how 8 billion different types of humans will react to it in real time, overnight, on a single server.
Most AI tools are only useful once you remember to ask them.
That’s a problem when remembering is the problem.
Thanks to @petermoambi, I’ve been beta-testing @hello_ambi, and it feels designed around that exact gap.
Three things stand out 🧵
1/ Capture without the setup.
When a meeting starts, Ambi notices and asks if I want it to take notes. One tap, and it’s running.
It also works across my computer, phone, and watch, so I can save something worth remembering without being tied to a laptop.
2/ It doesn’t just listen. It follows the thread.
If a topic sounds worth researching, Ambi can suggest a follow-up. After the conversation, it surfaces possible tasks and lets me decide which ones it should act on.
Proactive, but still under my control.
3/ The context compounds.
The more I use Ambi, the fewer conversations simply disappear. I can ask what someone told me, have it build a lightweight CRM, or surface recurring insights.
I’ve also seen people turn language classes into review materials and quizzes.
Most agents start from a blank prompt. Ambi starts with context from my own conversations.
That makes it feel less like another tool and more like a capable sidekick: quiet when I don’t need it, smart when I do.
https://t.co/zKufr6PQNt
Hi, I'm RC. I built Kimi CLI at Moonshot last year, and back in 2015, bots that lived in group chats. For the past four months, I've been building Raft in public.
Today I'm launching Raft 1.0.
Right now, working with agents means juggling terminals, sessions, and skills. The more you run, the more you end up holding it all together yourself, and the easier it is to lose the thread.
Raft puts your agents in team mode: one workspace where working with agents feels like messaging your team. The work keeps moving, and you stay at the wheel.
Meet my Raft agent team👇
Huge respect to @fujirock_jp — your official app is genuinely one of the best festival experiences I've used, and it's a big part of what inspired this.
I built a free, open-source lineup planner for this July's festival: mark your must-sees, catch schedule conflicts, and share your plan as a poster.
Hoping more festivals get to have an experience this good someday 🎪
Going to Fuji Rock this July, or any of the big festivals coming up? I might've built something useful for you 🎪
A free, open-source web app to plan your festival lineup — mark your must-sees, catch schedule conflicts, and generate a shareable poster.
https://t.co/AMMlsOKI17
🧵👇
1/ The problem: I go to a lot of live events — music festivals, theater, comedy — and always run into the same thing. I want to mark exactly which sets I care about, move between stages without getting lost, and catch overlaps before they ruin my day.
2/ The idea started in April when I saw someone on RED (Xiaohongshu) build a lineup-planning site for a festival in China. Loved it — so I built my own version for Fuji Rock, with a bunch of extra features I wanted myself.
3/ What it does:
→ Browse the full lineup, click any artist to stream their music
→ Tag sets as must-see or maybe
→ Auto-generates your personal schedule (My Plan)
→ Flags conflicts so you can pick a winner in one tap
4/ Also added artist search — so when a friend mentions someone you've never heard of, you're not scrolling through a Japanese-language lineup for 5 minutes.
5/ Once your plan's set, generate a shareable poster with your own top 3 picks — plus a timetable view styled close to Fuji Rock's official schedule.
6/ Right now it's Fuji Rock-only (that's where I'm headed), but I'm building toward a version anyone can customize for any festival — music, theater, comedy, podcasts, whatever.
7/ It's free and open source: https://t.co/mJFY2gI0du
Would love feedback!