Thank you @rabois for the rapid fire feedback and questions tonight. And thank you @eriktorenberg & team for hosting @beondeck in Miami. @btaleisnik @liamherbst29
The @BostonCollege Investment Committee (an LP and my beloved alma mater) asked for a few thoughts on what's happening in AI. I recorded a test run yesterday morning and then shared it with my partners, who encouraged me to share it more broadly... so here you go!
This is not a sales pitch, it's just a reflection on what we're seeing. And it wasn't intended to be shared, so please pardon the rough edges.
https://t.co/AuZqHMSSxD
Personal Agents Notes (Instinct / Grok Bots / ChatGPT Work):
- The three products to look at are Instinct, Grok Bots and ChatGPT Work. All three feel like the distillation of patterns established by OpenClaw earlier this year: persistent agents with cloud(ish) computers, browser access, cached credentials and recurring loops—plus a top-level orchestration agent with visibility into every other agent/thread that can report across them and dispatch work on the user’s behalf.
- Browser use is now good enough to do most web-based work, modulo CAPTCHAs / 2FA / occasional brittle flows. The interesting product variable is increasingly presumptuousness/resourcefulness: how much permission does the agent assume / how resourcefully does it recover from failure.
- Instinct and Grok are both very aggressive here. I had both shop and purchase things for me overnight and they were remarkably effective. Instinct couldn’t access one website, so it reset the password and completed the task .. resourceful!! Slightly insane!! Also a good example of why it may do things that ChatGPT Work probably won’t.
- Instinct has made every possible tradeoff toward a dedicated consumer product: iMessage as the primary interface, one continuous relationship and almost no visible machinery. This is probably the simplest mental model, but a single long-running thread creates real constraints—sufficiently ambitious work needs isolated context, tools and memory vs endless compaction. Definitely built on the very good lessons of Poke.
- Grok Bots feels more like an enterprise agent platform though it works well for prosumers as well. Named agents are probably more intuitive than threads for most knowledge workers, and features like connecting multiple accounts from the same service — personal Gmail, work Gmail, etc. — are very thoughtful. I’m less convinced that the agent group-chat metaphor will be intuitive but this may be more of a powertool while the average user simply chats with a single agent.
- ChatGPT Work is currently the least aggressive of the three, but probably has the best overall interface and balance of tradeoffs: an omni-agent front door, specialized threads underneath, full-duplex voice and access to both consumer + coding workflows. Its reluctance around credentials, payments and consequential actions feels like a compliance choice vs a technology constraint, I'll be curious if they close the gap as Grok/Instinct take off.
- Full-duplex voice makes this architecture much more natural because you can continuously steer the orchestrator while work happens asynchronously. Grok currently treats voice mostly as transcription; ChatGPT is much closer to the experience of actually managing many parallel loops through conversation.
- The cloud / local question roughly maps to knowledge work / coding. Knowledge work benefits enormously from an always-on cloud VM that can keep operating while you’re away. Coding still frequently depends on repositories, credentials and tools on localhost—although much more coding should already have moved into the cloud and probably will.
- The initial consumer wedge has been an open question for me... email + calendar are natural candidates but many consumers barely use a personal calendar and receive very little consequential email; productivity alone is unlikely to be the mass-market hook.
- The aha moment for me was seeing these agents shop for me .. /shopping may be the first loop that really sticks because it combines research + judgment + execution and produces an immediately legible outcome. It is one of the first use cases where the product clearly feels like labor vs software. But it is probably (hopefully!) the beginning of a broader set of consumer loops across family / finance / health / social / self-improvement—each of which requires some combination of information, motivation and follow-through.
- The transition we need to take consumers through is prompts → loops → sets of loops. The progression is natural in the enterprise because the loops are already explicit i.e. bug fixing, feature development, sales, support, procurement. Consumer loops are less formally defined but ultimately much broader and more consequential.
Arguably personal agents have far more implications for consumers + society than coding agents because they represent zero marginal cost work.
Every consumer can live a “fully hacked” life; every decision that requires information, motivation and follow-through can eventually be made and executed on their behalf. This means every software product can now be delivered as a service: family office for every family, concierge doctor for every patient, life coach for every person, etc.
The personal-agent market should be at least as large as coding agents and perhaps much larger. So I don’t think this is fundamentally a capability or demand question anymore. Browser use is good enough and the latent demand for labor is effectively unlimited. The open question is product design + distribution: which initial loop earns enough trust, context and permission to expand into the rest of someone’s life?
The steelman for a new consumer entrant is that Grok will split its attention between enterprise and consumer while ChatGPT Work remains embedded inside a much broader product. Dedicated consumer focus could produce a genuinely divergent product .. but the thing to underwrite is retention around a real loop, not the magic of the first session.
Excited to see where all this goes!!
aa + 5.6
As we’re seeing in case study after case study, it turns out that the amount of value that can be created between the AI model and the ultimate end-user workflow is far larger than many people assumed or realized.
Model capability is obviously doing a lot of heavy lifting in agentic products, but there’s still a lot more work to diffuse AI into the enterprise.
1. Getting agents to work well (and alongside people) in mission critical workflows tends to need to be represented differently depending on the business process. Sometimes it’s a chat experience. Other times it’s a background agent running in a deterministic workflow. And dozens of other variants. This is a mix of needing a harness that’s tuned to specific domains of work, but also making it show up in the right product experience.
2. Different workflows connect into entirely different enterprise systems and need access to very different data. Working with that data -whether it’s life sciences, financial, legal, etc.- requires contextual approaches, understanding of the data, having the right user experience for data interaction, and more.
3. The need for domain-specific change management remains critical in most verticals. The way you talk and implement technology at a bank is very different from a law firm. Having the right talent with a singular mission ends up being extremely useful for something as complex as process reengineering.
4. The ability to work with a variety of models means you can tune the workflows to different cost and performance levels. And you can eventually post train models for specific tasks to tailor the outcomes and eke out gains that aren’t coming otherwise in frontier models.
5. Evals! AI is basically not useful if it can’t be evaluated. Domain-specific evals that let you dramatically improve the performance of your harness for specific workflows just has a crazy long tail given how many tasks there are in the economy. Nearly impossible for one system to be tuned for all of them.
6. Lots of verticals and domains require pricing models that reflect relevant abstractions on top of tokens alone. The ability to price in ways that work for your industry’s consumption model ends up mattering in a variety of spaces.
This just touches on some of the things that go into the applied AI layer. But it all adds up to being a huge surface area for being able to sustainably innovate and differentiate.
Hard things I learned in scaling a manufacturing business that I wasn’t fully prepared for:
- as the company gets bigger, using slack to communicate across all shifts and factories in different states sucks. You gotta try harder than slack. Things can get taken the wrong way (gotta use lots of emojis) and important messages get lost. Phone calls and travel are so much better but you gotta put in the effort
- with 600 employees you don’t know everyone, but you still feel like they are your kids in a way. When someone loses a family member, or they have a sick kid, or they get into a motorcycle accident, I worry about them like they are my own kid. I can’t imagine how this is going to work at 6,000 employees.
- entropy is real. We are fighting laws of thermodynamics lol. Everything wants to dissolve into disorder and chaos. Gotta fight it 24/7
- it’s easy to make negative assumptions. I see front office teams grow and I’m thinking “really? It’s a full time job just to do xyz??!” But then I sit with them and realize at scale, yes…it’s a full time job to do things that took us 20 minutes a week in the old days
- many vendors and salespeople don’t give you the time of day when you’re small or starting out. Then you get to a certain size and they are begging for your business, offering discounts, and giving you access to crazy deals. If you survive the valley of death it’s nice on the other side
- safety and compliance becomes a big deal. When it’s just 5 of you, safety is implied and everyone kinda gets it because everyone has “done it” before. With a larger population, you have to worry about so much more. Some people are brand new to manufacturing and don’t have the instincts that come with working with machines for years and years
- the larger the company grows, the more you have to lose. hundreds of families count on me for a paycheck, so the stakes are really high. Easy to lose sleep when things go sideways for a little bit.
- when you first start, things seem to cost a couple thousand $. Then after a while it’s $10k. Then $100k. Then you get comfortable with $1M or more and your concept of money is wildly skewed. Every quote you get is some batshit crazy amount of money but apparently that’s how much things cost
- mistakes can cost many millions of dollars at scale. You can still take risks, but the downside starts to get bigger
- you can’t be everywhere all the time, so you gotta have great people that would make the same decisions you would make if you’re not in the room. Or ideally they make an even better decision. Great leaders allow companies to grow. Build them early
this makes me extremely bullish on g. sergey should take even more control of the co & delegate everything that isn’t core to sundar. more importantly he should also be the public face of google ai cuz he is extremely articulate & very likeable too.
google’s biggest constraint is org velocity (bureaucracy) + external comms / messaging. sergey is the only person that can bulldoze through both of these by being authentic, relatable, & honest.
also would not be too shocked if sergey starts posting on x.
Elon Musk reveals the ego-to-ability ratio that predicts whether someone will fail
"You do whatever it takes to succeed, and just always be smashing your ego. Internalize responsibility. A major failure mode is when your ego-to-ability ratio is greater than one"
"If your ego-to-ability ratio gets too high, you break the feedback loop to reality. In AI terms, you break your RL loop. You want a strong RL loop, which means internalizing responsibility and minimizing ego"
"That's why I prefer the term engineering as opposed to research, and I don't want to call xAI a lab, I just want it to be a company. Whatever the simplest, lowest-ego terms are, those are generally a good way to go. You want to close the loop on reality hard"
Interesting take. The challenges today:
1. Most enterprises don't know how to make AI game changingly effective. As they embark on the journey, the use cases being addressed are "80%" single shot, semi deterministic use cases with multiple guardrails or humans in the middle. We are a little ways away from mass adoption of custom models.
2. More complex cases which require any multi agent orchestration and context retention are beginning to be conceived and tested. These cases will make model portability harder, requiring new evals and harnesses. One will have to commit to one structure and also commit to constantly updating and retraining your model..
3. CIOs and CEOs aren't sure if the ultimate architecture is single stack, multi model - interoperable orchestration and context/harness/eval, or a custom model. Uncertainty causes slowdown on longer term decisions, which in this case is perhaps right.
4. Custom and Opensource come with the need to deploy on either your own GPUs or public cloud. "Interesting fact - if token prices fall as I hope - it will be cheaper to run frontier LLMs than open source on your own GPUs"
5. Generally horizontal solutions that can serve tens of thousands of customers make more money than vertical custom solutions, but maybe this time it's different?
Even if we solve the model conondrum, the enterprises need to redefine workflows, collect more training data on each use case and rebuild the application in a simple UI flow, not everything will be done in a conversational window.
But time will tell, this will continue to be a space to watch, lots of minds at work to solve this.
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor.
It’s happening faster than we thought, and the implications deserve greater attention. https://t.co/OVVPJO7VQx
Starting to hire and retrain for new agent engineering roles for *internal* functions to help get more powerful agents working well on critical business processes. I expect this type of role to be a very big deal over time at Box and other companies.
It looks something like an internal FDE, whose job it is to wire up internal systems and get agents working with them effectively. The person will be extremely technical and capable of building secure, governed agents for internal workflows that connect to business systems (like Box, Salesforce, Workday, etc.), and codify workflows in skills.
In some cases this person may understand the business process well enough to do it fully, but in most cases I expect them to work with the business directly in an embedded fashion. Ironically, that may introduce another new role on the business side that is more akin to agent product management for internal processes. The key is that you need technical + process people that can span multiple teams or functions in an organization. It’s not about brining automation to a job, but bringing automation to a process.
This is going to be a very big trend in most companies going forward. Fun to watch the early innings of what this will look like.
Google DeepMind’s real-time video AI doctor is here.
They just introduced AI co-clinician, a triadic care system built to work under a doctor’s supervision during patient care.
The system is built to retrieve clinical-grade evidence, verify it, and in patient-facing simulations use a dual-agent setup where one module talks while another watches for boundary violations.
It also beat other frontier models on open-ended drug questions, because real medicine arrives as messy patient cases, not multiple-choice exams.
DeepMind evaluated it against the failure modes clinicians actually care about: saying the wrong thing, or failing to surface the crucial thing.
In 98 realistic primary care evidence queries, physicians preferred the co-clinician to leading evidence-synthesis tools, and the system logged zero critical errors in 97 cases under their NOHARM-style evaluation.