I'm genuinely curious about this. On one hand, it is a great idea (s/o @monogram_ai) and improves the UX a lot, on the other, significantly more of the effort, context, reasoning power is spent on building the UI and tooling which takes away from the problem at hand.
ChatGPT just got a lot more visual—introducing Intelligent UI.
Intelligent UI allows ChatGPT to answer quickly with fully interactive user interfaces that make everyday answers more visual.
It makes learning complex topics easier, and it quickly creates tools to solve a task right in the moment.
Oh, and also...GPT-6 is coming to ChatGPT. For everyone.
@SteveBrecher@eric_seufert I don't claim they do or they will. Meta does not need to share conversation history or VM data to mark one's interests or purchase intent. I guess it would be terribly inefficient too, when you have a capable llm already handling it.
there are so many everyday demos that could make ai, personal agents, & agents in general feel insanely powerful to normal ppl & yet ai companies keep choosing things like booking flights, planning travel, or god forbid planning a god damn wedding.
these are not everyday problems. they’re relatively rare events & worse for a lot of ppl they’re actually part of the joys of life. the best demo should make someone feel pain they already are familiar with disappearing. almost every company gets this wrong. they optimize for spectacle instead of relief which is fine in some cases but it doesn’t last.
one of my favorite historical examples is a product called head on.
the ad showed someone with a headache, then showed them applying the product directly to their forehead while repeating:
“head on. apply directly to the forehead.”
beautifully simple. you instantly understood the product, the problem, & the value. ai companies should be doing the same thing.
show me the annoying thing i deal with every single day then make it disappear. it’s not sexy, i guarantee it will work. when jobs demo’ed the iphone he picked universal simple problems that the iphone did better than anything else on the fucking planet.
anyway, thanks for coming to my ted talk.
@ashray_malhotra You can always use an assistant that respects your privacy, only acts on what you say, doesn't connect to your email, and uses its own identity to get things done. Like @heysamwise
After years of working on Fat Goblins, I can finally properly announce it:
Fat Goblins releases November 12th!!!
It’s surreal to finally have a date. Thank you to everyone who’s played, backed, shared, or followed the game along the way.
See you in 16v16 on November 12th!
I don't really understand why agent to agent communication should be a thing. I want my agents to be as far away from other agents as possible, the more deterministic a service is, the better.
What do people think is the most likely future for consumers:
1. Everyone uses one ai agent for everything
2. Every major vertical has an agent and we use multiple agents daily
3. We use one agent and it calls out to 3rd party agents when needed behind the scene (variant on 1)
Be it coding agents or personal assistants, I think isolated virtual machines will be the building blocks of agentic workloads. Sharing your personal computer with an agent will be considered something we had to endure.
I mean yes, sure, but you also have power users like @erenyanik who tends to update his calendar every 30 seconds or so, resulting in even more wakes with this approach.
Also compute tends to be overstated almost by an order of magnitude in recent posts of this nature imho.
https://t.co/TrGIIuWV3m is $0.036/hour for 4 vCPU, 8 GB RAM, 50 GB.
This is p much back of the envelope but AWS's i7i.xlarge has 4 vCPUs and 32 GiB with virtualization capability, on-demand pricing is $0.3775 rn. The workload is sparse cpu-wise, mostly spent waiting on the model, and memory bound, meaning, you can fit in probably about 15 Instinct VMs at 2 vCPU and 2 GiB memory, resulting in $0.0251/hour. This is BEFORE any savings from reserved instances, or moving the workload to another provider like Hetzner, which should drive the cost further down.
It then becomes a utilisation problem, whether you can saturate your reserved capacity - which I assume Instinct has no issues with.
Instinct just got valued at $10 billion, and everybody in sf is building their own
i wrote up a guide on how to not burn money doing it https://t.co/3Ej5XUhrtM
If you live on twitter, you’d think agentic commerce is mostly about booking the sexiest restaurants in nyc and int'l flights
The actual killer use case that leads to agentic adoption: a busy parent uploading three back-to-school calendars and saying: figure out everything i need to buy and when/where I need to show up
WAYYY less luxury concierge, more making everyday life manageable
^ this is who we're building for @stripe
s/o @rcsenel 👐
If you live on twitter, you’d think agentic commerce is mostly about booking the sexiest restaurants in nyc and int'l flights
The actual killer use case that leads to agentic adoption: a busy parent uploading three back-to-school calendars and saying: figure out everything i need to buy and when/where I need to show up
WAYYY less luxury concierge, more making everyday life manageable
^ this is who we're building for @stripe
s/o @rcsenel 👐
@SebAaltonen@XorDev Agreed, but I guess his point is how many Opus releases before they are good at architecture. I used to feel LLMs converged on the generalised median, resulting in mediocre architecture but I think with the expansion of parameter space they are getting better every release.
@auchenberg@BaradAgent@heysamwise uses this infrastructure slightly modified, runs chromium inside these VMs. User data is restricted to chrome profile so hydration is fast and per-browser task. Dehydration is immediate after task completion.
@auchenberg We are building something relevant with @BaradAgent it spins up a new VM per session, can hydrate in seconds and dehydrate the VM after 10 minutes of inactivity. You send a new prompt to an old session, rehydrated within sessions. Disk persists, memory and processes do not.
I think we got it wrong.
I don't like the idea of AI snooping through my emails, my WhatsApp history or my documents folder to get things done. I don't want AI being smart about what to do, trying to push me into doing things it deems important. I want my AI to be smart about how to do what I ask of it.
I don't like the idea of ambient intelligence being the future either. It removes intent from interactions. Look at the algorithms taking over social media feeds. I used to have a feed filtered via my intentions: people I followed, people I am friends with, people I chose to give attention to. Now the feed is ambient. I have very little agency over what I see. The algorithm decides which shorts I should be watching, and more often than not, it is the same viral pick of the day.
AI should not take over my identity either. An agent can do a great deal more with my identity than without it, and many tasks still require a human one: accounts, payments, purchases, services. But, it takes away the authenticity and trust we've built with each other. When I message my cofounder Eren, I trust that it is Eren reading and responding to my messages. If an agent responds on behalf of Eren, pretending to be him, it erodes that trust. AI agents should always have their own identities, even and especially when acting on behalf of a principal. I don't mind interacting with agents as long as I know I'm interacting with an agent.
This identity distinction matters for AI itself too. Is AI a tool, or an entity, actor, principal? I don't think we're really treating AI as a tool anymore. We expect it to have judgement, refuse instructions, and follow its own rules. Claude wears its own values pretty proudly, and I don't have a problem with that. My point is: if there is another actor in the loop with its own judgement, it should act under its own identity, not mine.
These thoughts pushed me to start working on two products: Samwise and Barad.
Samwise (@heysamwise) is a personal agent with its own identity. I can forward an email chain to Samwise and ask it to take over. It will reply to emails, chase people, manage calendars, send follow-ups and let me know when it needs my input. But it does all of this as Samwise, not as me, and only when I instruct it to.
It will not impersonate you. Samwise will always use its own name, its own email address, and its own telephone number. It will always disclose who it is acting on behalf of.
Samwise called restaurants to make bookings for their users, to cancel their monthly subscriptions, and to book them a parking space at the airport. But it did this all by being itself, Samwise, an agent who works on behalf of a principal. The person interacting with Samwise knows they are interacting with an agent, not me, not an agent identifying as me. They might bypass Samwise and address me, or ignore it because it's an AI. That's fine.
Barad (@BaradAgent) is a remote coding agent. It is a BYOK remote harness. You enter a prompt to spin up a durable virtual machine. It goes to sleep after a period of inactivity. You send another prompt, and it comes alive again.
This has allowed me to set my agents loose. I was very protective of my computer before. I was checking permissions, using --dangerously-skip-permissions when I craved adrenaline, or spinning up a VPS if I felt the need. I wasn't allowing my agents to take over my browser or roam around my hard drive. Well, because it is mine. It has my child's pictures, medical records, and years of my digital life.
With Barad, my agents have their own computers. They are free to use the browser, record a video, take screenshots, and install whatever they need without me worrying about what they are doing to my machine. If I don't like it, I trash the VM and move on. My computer is still my computer.
I gave my agents more freedom by separating them further from my own machine.
I think this is the better answer: give AI its own identity and tools, and distance it from mine. I want my email to be mine. I want my computer to be mine. I want my browser history to be mine.
These are very early days, both for the change we're living through and for these products. I would genuinely love for you to try either one and share what you think.