My guess is the split is less about attitude than about how your work arrives. The 25% mostly have jobs that come as a queue of discrete tasks, which is exactly what these tools eat. Household life doesn't come that way. Nobody wakes up thinking they must ask a machine about Thursday's dinner; they discover at 6pm that nobody has thought about it. The tools will reach the other 75% when they stop waiting for a question and turn up with Thursday already sorted, asking only whether that's alright.
Just wild divergence among friends/acquaintances right now in AI use, maybe 25% are like "this is making me massively more productive" and 75% are only using it at the margins. Very little middle ground. (I'm sure the 25% is very high relative to the general population.)
@omooretweets@instinct@Muse@tomo@bot Lovely list. The most important agent in it is still you, at 10pm, noticing that the one tracking your sleep has no idea another has booked dinner at 9. And none of the seven seems aware you live with anyone. Seven slices of a household still leaves someone doing the joining-up.
This is the bit that matters more than the model. Nobody wants to read their week as a paragraph; they want to see it laid out and then just say "move Thursday". Conversation to change things, a screen to see them. The interesting step after that is the screen turning up before you've asked.
I suspect half of the "still waiting" list is one category in disguise. Travel, finance, food, health: in most homes these aren't separate apps, they're the same Sunday-night conversation about what the week can afford in money, time and sleep.
Vertical agents keep losing because each one only sees its own slice. The trip planner doesn't know it's a heavy work week. The budgeting app doesn't know the kids need new shoes. The winner may be whoever plans for the household rather than the category.
@pitdesi Personal assistant fatigue is a very modern problem: we've hired six assistants and now manage all of them. The one that wins won't have the best pushes. It'll be the one you realise you've stopped opening, because the week just went fine.
Right problem. "What are we eating?" gets asked roughly 365 times a year in our house, and the answer has never once been a recipe.
But look at the setup step: favourite meals, cooking time, foods to avoid, optional budget targets. That's a questionnaire, and households are famously bad at filling those in honestly. Nobody writes "we say we'll cook Thursday, then order a curry."
A good meal agent learns that from the receipts. What actually got bought, what got thrown away, which nights the calendar ate the cooking time, whether everyone slept badly after the late takeaways. Then it drafts the week and you correct it.
The budget shouldn't be a number you have to set either. It should be how the month is actually going.
If you’re going to build a personal agent, a Meal & Grocery Agent should be near the top of the list.
Not to throw random recipes at you. To take the weekly “what are we eating?” problem and turn it into an actual system.
Think about what that takes off your plate:
› Fewer last-minute dinner decisions
› Meals that actually fit your schedule
› Better use of what’s already in the fridge, freezer, and pantry
› One organized grocery list built around the plan
› Easier adjustments when the week changes
› A simple way to keep leftovers and ingredients from disappearing from the plan
Here’s the blueprint for building one.
› Plan meals around the nights you actually have time to cook
› Work around your household’s likes, dislikes, and usual meals
› Use what you already have in the pantry, fridge, or freezer
› Combine what’s missing into one grocery list
› Rework the plan when your schedule changes
› Keep track of leftovers and ingredients you still want to use
Then you make it yours.
› How many people you’re feeding
› Favorite meals
› Cooking time
› Equipment you actually use
› Foods you want to avoid
› How often you’re okay repeating meals
› Optional budget targets
The blueprint below gives you a starting structure for the routines, memory, tools, rules, automations, and outputs behind the agent.
Build the structure once. Then shape it around how your household actually eats.
Good to see an agent given a screen as well as a mouth. Chat is fine for changing things and hopeless for seeing them; nobody wants their week read aloud. The real test of proactive is the quality of the interruptions. One Sunday nudge that gets the week right beats twenty clever ideas on a Tuesday.
I think the empty text box is the biggest design mistake in personal agents.
It assumes I know what to ask. Most weeks I don't. I find out on Thursday that nobody booked the dentist and the car is needed in two places at once.
Anyone who runs a household well doesn't wait to be asked. They sit down on Sunday and draft the week. An agent should do the same, unprompted:
> Sunday evening, a draft of next week just shows up
> who's on pickup, what's for dinner, what gets booked and paid
> planned around how everyone's sleeping and what we can actually afford
> I fix what's wrong and it gets on with the rest
Editing a draft takes five minutes. Writing one from a blank box takes an hour, which is why nobody does it and why most agents sit unused after the first week.
It's also why I don't think chat alone wins. You talk to it to change things, but you need to see the week laid out to know what's wrong. The good ones will be half conversation, half screen.
And the draft is only the start. Every week it gets right is a week I check less. The end state is that the week just happens and I only hear about what changed.
Exactly this. Booking the table is the easy 5%. The other 95% is knowing it's the third takeaway night this week, the kids have swimming Thursday and money's tight after the car service, then planning the week around all of it before anyone asks. One-off tasks are a demo. Running the household is the job.
Good that someone's measuring this. But booking a flight or moving a meeting is a task, not a life. The benchmark I'd want runs for three months on a real household. Did they sleep better, spend less on stuff they didn't care about, and get hours back? How many decisions did nobody have to make? Task completion is the floor. Outcomes are the score.
Which personal agent is the best life assistant?
Today we're releasing the micro1 PersonalAgentBench to help answer that, and we're paying real personal agent users to contribute to the benchmark.
Personal agents can now book your flights, answer emails, and move your meetings around. That’s great, until one messages the wrong person or makes a purchase without confirming.
Capability is not the same as trust. So we put four personal agents to the test: Instinct, Muse, Grok Bot, and Gemini Spark. We found that even when agents found the right information, they still often miss what matters, overshare, or make things up.
To find out for yourself how your agent compares, you can join our benchmark. We’re paying the first 100 users to run our prompts and share their results.
Proactive isn't the problem. Scope is. An agent that can see your bank balance has to know which room it's standing in. Money context belongs to the household, not the company Slack. Any agent with real access needs a hard line on who's allowed to hear what before it gets to be clever.
@aliansarinik Capability without trust is just a faster way to book the wrong thing. I've been saying since Dec 2024 that weekly planning only works if bookings happen zero-touch with minimal involvement — and that means the agent has to earn the right to act, not just find the flight.
One-off bookings are the demo. The real job is the recurring mental load — calendar, groceries, birthdays, bills, dentist, what we're eating this week. I've been saying since Jan 2024 that a self-driving calendar was the point. Zero-touch is the end state, not another chore list with a chatbot on top.
@chiefaioffice I don’t think he’s raising $7tn to make chips. I think he’s raising $7tn to buy up all the companies that would comprise an ecosystem for full AI automation for people’s daily lives. A grocery store, travel agent, bank (with deposits), streaming platform, compute provider, Uber..
Said this in Feb 2024 when everyone was talking chips.
The scarce thing was never the model. It was the real-world gatekeepers — grocery, travel, bank, calendar — that make zero-touch life automation possible.
That thesis aged well.
@BrennanWoodruff@GaryMarcus And he’s very good at stoking the soap opera crowd. But nevertheless, someone is going to need that ecosystem of real-world gatekeepers to deliver a truly transformational AI experience of zero-touch life automation
@BrennanWoodruff@GaryMarcus And he’s very good at stoking the soap opera crowd. But nevertheless, someone is going to need that ecosystem of real-world gatekeepers to deliver a truly transformational AI experience of zero-touch life automation