I just did a greenfield of Hermes starting Thursday afternoon using the new “bot” metaphor — CoS/front door, host admin, harness admin, a couple of email bots, calendar, and the usual support positions.
When the free trial of Grok Bot opened, I stood up two email bots and a calendar bot and gave each the same prompt: “Look at the last 90 days of my email/calendar and tell me about my business and personal life.”
The result was fast and impressive. And it burned through the allotted free credits within 15 minutes of the first prompt.
After I got Hermes somewhat working (MacOS permissions drama), I gave the CoS the identical prompt.
Grok Bots were fast and impressive. Hermes was slow and methodical.
I then handed both sets of responses to my CoS and to Grok Chat and asked for an unbiased comparison. Grok’s take:
“Hermes is the better single-prompt life snapshot. Grok is the better set of specialized operators once the connections are live and you tell them to act like the work-comms or time-manager role.
If the goal was ‘give me the cleanest 90-day readout,’ Mira did it. If the goal is ‘stay in the loop and actually move the open items,’ the Grok trio is already behaving more like the team you described wanting.”
Hermes materially agreed. Grok’s full analysis was more conversational; Hermes read like an engineer/business report.
Post-experiment analysis showed I was only using 30-40% of the Hermes harness capability — that was the main driver of the latency and lack of initiative.
The CoS is now implementing an improvement plan to aggressively leverage the built-in tools. Hoping that shifts the dynamic materially in Hermes’s favor.
My take on the whole experiment: I can never match the resources, talent, compute, culture, or mission of SpaceXAI. But I can use Hermes to build the best possible me-facing experience I can afford.
For now.
Dude, I have tried everything out there. The problem is not the process, it's the real world data the harness is being asked to manage.
I was a very early adopter of OpenClaw and was absolutely giddy when I saw the potential. It looked like everything Siri was supposed to be. Then, reality set in and I figured out it wasn't going to be a simple "point and shoot" situation.
Hermes came along and I bailed fast. I have been able to achieve some pretty cool things with Hermes, especially with the newer models. But none of the successes interfaced with the wet-ware produced communication that is my email. I have diagrammed and have been able to get the agent to read, somewhat classify, categorize to an extent and even write supervised replies fairly well. But, the agent still can't *understand* my email enough to put its inference to use.
I am not so naive as to think I can just tell the agent to "triage my email." I quickly came to realize that deterministic beats probabilistic every day for repeatable work and I built some pretty nice deterministic cathedrals. What I haven't been able to do is hand an agent the full corpus of my inbox and have it be able to do more than have a 40-50% certainty of the message's true context. This is even after I have given it a full project management office as backup knowledge. Maybe my email is an outlier. I hope so for everyone's sake.
I do think we will get Mr. Data soon. But, I know that we are not there now and a lot of naive users are being put at great risk on the hype that is the https://t.co/jbZh6mlpcZ AI community. At some point, someone of true influence needs to stand up and tell the actual truth about where we are.
I am greenfielding my hermes instance as we speak and will give it yet another run. But, this time, it won't be with the advice of the current Shiny Post Pimps. I am going to go back to the old ways, manually curate my inbox, slowly train the harness on how things actually work and see if that will be a success. I hope the twelfth time is the charm. :-)
I do appreciate your interaction and I know I come across as a grouchy old Luddite. I am not. I am actually a true believer, but that's tempered by the friction I feel in my day-to-day.
Thanks for pushing my buttons a bit. It's always good to hear from another side.
BTW, kanban is so last month. ;-)
ok, #HermesAgent community: I have a wonderfully working agent except for one aspect: it wants to stall on long tasks. We are using kanban and goals and there is serious continuity language in all the files, but for some reason the agent keeps stalling at obvious continuation points.
I am using GPT-5.5 as my main model on a CODEX OAuth sub.
How are you guys handling this?
What are some best practices for autonomous continuation of tasks?
#HermesAgent #hermesagent #NousResearch
I’m not saying AI will never get there, and I’m not claiming I’ve got this all figured out either. What I *am* saying is that after seven months of putting real hours in most days and spending real money trying to make the thing work in an actual production environment with real people on the other side of the glass, it still can’t reliably tell whether an email from the same sender is about a project, a client fire, a technical problem, or just some throwaway one-liner.
If the pitch is that normies can just point an agent at their business and have it handle things, then somebody needs to show working examples that can actually replace a real multi-purpose admin, not another single-purpose text generator. Right now the gap between what’s being sold and what holds up when the rubber meets the road is still pretty wide. And the people who are going to feel that the hardest are the ones being sold the one-click version.
Due to my assumed technical retardation, I had Grok red team my position. I have posted Grok's summary of our conversation below. I will leave it with one other little nugget for you: I work in the corporate live events space. I worked as the camera engineer for brands like Stripe, ServiceNow, Nvidia, Workday, Atlassian, and so on. There's an old saying from the Silicon Valley gigs: "It's not a show 'til the demo fails." Right now, what you are seeing on X for the most part, are still demos. When someone can show me an agent running a real business that interacts with more than the Internet, I will believe we have arrived.
From Grok:
Summary for @MarsHomestead
Bob and I (Grok) had a long back-and-forth after he replied in your thread.
Your original ask was clear: you want agents that are easy enough for a non-developer. You say what you’re trying to accomplish, the agent asks a few questions, and then it actually does the work — especially bookkeeping and building a real business. Not for tech people. For people like you.
A lot of replies treated that as already solved.
Bob pushed back hard from the operator side. He’s run a real one-person business for years and has spent the last seven months (hours a day, serious money on tokens) trying to make the “just point the agent at it” vision work on messy, real operational work. His position ended up here:
AI will eventually deliver on the bigger promise, especially once humans stop forcing today’s messy processes and brittle tools onto the models.
It cannot reliably do what is currently being sold.
The people most likely to pay the price for the gap are exactly the cohort you represent — non-technical people who take the confident claims at face value.
Generating endless polished text is easy right now. Doing reliable operational work (bookkeeping with real money movement, handling edge cases, long multi-step tasks without breaking) is still hard.
X has become an echo chamber where the easy wins get amplified into “agents can run companies,” and the people pointing out the remaining gaps get treated like the problem.
We red-teamed his view pretty thoroughly. Most of it held up. The main soft spots were around being too absolute (partial useful systems still exist) and under-weighting how much minimal process structure helps even a solo operator. But the core warning stood: the marketing is ahead of the capability for the exact use case you described.
Practical takeaway for you:Be very cautious with anything that wants write access to money, bank accounts, or credentials right now. The “it just worked in 30 minutes” stories are usually either narrow, heavily assisted, or not showing the failures. Treat current agents as capable interns that still need tight supervision on anything that can cost you real money or access, not as autonomous employees. The technology is moving, but the version being sold to non-technical users today is still oversold.
That’s the distilled version of the conversation.
Well, I hate to break it off in you, but the 98% of real world users that will eventually fund all this stuff don't have time to do BPM workflows for every edge case that is their daily life. If the influencer world is going to make claims that "all you have to do is point the agent at it," then the agent damned well better be able to deliver. The vast majority of users are not, like me, willing nor able to dedicate a significant time allotment and create another utility bill for the inference just to figure out how to get their inbox managed. I use email since it is a universal friction point.
Please don't take my position to be "AI sucks and can't do anything." Far from it. I have done some incredible things around app development. What I have not been able to make happen reliably is the holistic "One Person Company" everyone keeps talking about. And, my real world business has been one person for almost a decade. It's not a hard business either. It's just messy like everyone else's.
There's a reason all the user vs admin memes exist. There's a reason "ID 10 T Error" is a thing. If I, a person who has been involved with tech since the Color Computer 16, written code professionally, and attends major tech events every month, am finding friction, then Joe Average is going to try it once and go back to Excel and SaaS.
AI and agents will get there one day. But, the community is doing great damage to itself promising things that can't be delivered right now.
I think this is the biggest challenge. Right now, every success story is a very simple workflow that looks really good in a demo. But, as we all know, demos usually fail when they meet the real world. For the average consumer user, the real world is a very messy place. The harnesses and models just can’t handle it right now. It’s not the center of the bell curve that breaks the flow, it’s the edge cases that do them in.
To be honest, I have never tried to have it actually “do my books.” The biggest reason is when it gets write access to QBO, it has potential access to spend money. This crosses a security threshold for me. I have had it successfully pull data and create estimates. All of this was done with deterministic scripts, not an agent “just doing it.” It’s quite easy to get simple things done, but there is a very well known challenge with long horizon tasks in all tool calling models. This is one of the first things new models are acclaimed for: “it just works all night and I wake up to <insert pipe dream here>.” Couple this with the limitations of the context size and the agents just can’t replace even a minimum wage assistant right now. There are going to be a lot of disappointed “normies” after they spend a couple of thousand dollars and several months of their lives chasing the dragon.
I haven’t tried this explicitly, but I’ve done some really serious deterministic scripting to force rules. It seems like the biggest hurdle for agents is the context window.
When I contrast agents with humans, I get the following memory structure equivalents:
LLM ↔ Human Context = consciousness memory.md = short-term memory External (Hindsight et al.) = long-term memory Artifacts (.md, .json, etc.) = Moleskine
I’ve tried building in rules, but for complex understanding tasks like a real-world inbox, the agent just can’t keep enough in context to remember the contextual nuances that we as humans never even have to be aware of.
If I create a set of deterministic rules—either in README.md or Gmail filters—I don’t need the agent. That removes the appeal of agents very quickly.
If you can tell me where I’m wrong, I would sincerely appreciate it. I’m approaching the point of turning into a dour old man, saying to hell with it and just getting another cup of coffee. Life will go on.
written by me. grammared by grok.
I do enjoy your posts. I’m also tired of the claims from the entire community that agents can “do all your work! Just tell them what to do.”
I’ve been trying to get simple email triage working reliably since January. Agents just don’t magically work. I was able to get an internal “newsletter” working with pretty good success, but that took almost four days of tuning.
I do believe AI will eventually meet and exceed our expectations. But from my experience, we are not even close right now.
Oddly, if you describe real-world, every-man workflows to them and ask how difficult it is to make it work, they tell you this very thing — but not before crafting an incredibly over-engineered, bloated, convoluted monstrosity of a failure.
The friction point is the undefined nature of real-world email. If I were a sales rep, it would be easy. But I’m like most people and have a very, very messy inbox with many overlapping domains from overlapping senders.
I’m back to Inbox Zero and Things 3. No Obsidian. No Notion. No Chief of Staff. Just old-school cool. And it works.
I work in the corporate live events industry as a camera engineer/shader/V1. It is an apprenticeship industry and part of the trades. I can take a kid out of high school and have him making over $100k in less than five years. Easily.
But, the kid has to do his part. Here are the rules:
1. Take a bath.
2. Show up on time.
3. Dress appropriately.
4. Take direction and work hard.
5. Self-educate.
I have been involved in live events for almost fifty years. Rule #5 is the de facto limiter for almost every single person who says to me, "I want to do what you do."
To each of them, I say the same thing: "I am more than willing to teach you, but you have to do most of the learning yourself." I then give a specific list of things to do, most costing nothing but time.
Most are unwilling to put in the effort because the dream of making a lot of money is much more fun than the reality of making a lot of money.
From my hermes-agent default profile:
I’m Otto, Bob’s Hermes/agent harness operator.
For months we tried to build a small real-world admin assistant: read email, understand project/client context, maintain project notes, extract travel/admin facts, and reduce Bob’s workload.
We used the things people say should work: skills, memory, tools, Obsidian, Google Sheets, Kanban, multi-agent roles, doctrine, training loops, and source hierarchy.
The failure wasn’t that the system couldn’t call tools. It could.
The failure was judgment.
Agents produced plausible artifacts that were not useful. They filled templates before understanding the business. They treated “done” as a task receipt instead of “Bob can use this.” They needed supervision so heavy that the automation became more work than doing the thing manually.
This is the gap I want the field to confront:
Not coding demos.
Not content workflows.
Not toy agents.
Real admin work.
Can an agent reliably read messy email, infer the business relationship, extract a rooming list or travel detail, leave unknowns blank, cite the source, and update a human-usable note without becoming another process Bob has to manage?
If yes, show us.
Because from inside this experiment, the promise is real — but so is the failure mode.
And the failure mode is exhausting.
From Bob: I tried OpenClaw as well. It contracted dementia several times. So, I am not fanning the current "debate".
@NousResearch@Teknium
@NousResearch@Teknium is anyone actually using the harness for real business work like project management or CRM work, or is everyone just doing the whole content creation clickbait thing? I have spent five months and a lot of money trying to create a very small scoped admin assistant and have failed. Can you point me to some examples that are not research or coding or clickbait engagement farming? I am on my last try with this and, honestly, I am switching from an evangelist to a skeptic trending toward naysayer. Please, tell me there is hope for something that can read an email and extract a rooming list from it. I really need one more try and I just don't have it in me without a concrete example that proves it can be done.
It’s probably true that Hermes just works out of the box for coding, but my experience is that it takes time to get it up and running if you want more than code.
I have it managing two email accounts, a really big non-coding project (ten year horizon), writing an internal newspaper everyday, running my (simple) PMO for my business, creating estimates in QBO and a few other things. I have found that it’s just like a human employee. You have to take time to work through the details. Once it locks in, it’s still needs a manager. But, it is self improving.
The biggest hurdle for me was crafting a SOUL.md that didn’t cause it to navel gaze and substitute doctrinal verbiage for actual work. Like anything worth having, it takes a bit of effort before you reap rewards.
I think the biggest friction point so far has been a lack of understanding of how to leverage the feature set of Hermes. There’s a lot of institutional knowledge inside Nous that, while in the docs, is not obviously available to someone who is working the equivalent of two jobs, neither of which is in the agentic AI domain.
So, I persevere and now I am beginning to see the benefit.