Somewhere right now @poteto and a bunch of very serious people at xAI are reading my @bot chat history…
where I alternate between swearing at Grok Bot, trying to humanize it, arguing with it, apologizing to it, and saying completely unhinged shit 😬
One more bit of context: maybe I simply expected too much from Grok Bot.
I’ve been using OpenClaw agents for ~6 months already, and by now they handle most of my routine work. Simple automations aren’t really the problem for me either — if a task can be solved with a script that pulls data from a few sources and combines it, I can usually build that pretty quickly.
What I hoped Grok Bot would give me was something beyond that: an agent I could give a large amount of context to, that could actually retain and use it, understand how I work, and take ownership of messier tasks that aren’t easily reducible to a deterministic script.
Maybe I got overexcited after reading all the enthusiastic posts on X 😅
What surprised me is how little I can actually inspect or control: I can’t really choose the model, reasoning level, see/manage its filesystem or memory in any meaningful way, etc. So when something goes wrong, I also have very little ability to understand why.
Also… I’ll admit I was slightly embarrassed to send you the chat ID because there are definitely moments in there where I get frustrated and argue with the bot 😂
Hopefully that isn’t the reason it started performing worse…)
Poteto poteto poteto
@poteto@bot — maybe I’m doing something wrong, but I genuinely need some advice here.
I use a lot of different agent systems — Hermes, OpenClaw agents, Grok itself, etc. I paid $300 specifically to get access to Grok Bot. And yes, I know it’s much cheaper to access now, which honestly makes this even more frustrating.
So far, Grok Bot has been the worst agent experience I’ve had.
Even a ridiculously basic task — preparing my evening report — keeps failing.
The workflow is documented in detail. I recorded TWO ~10-minute videos showing exactly how to do it. The bot understands the task, can explain it back to me, schedules the routine for the correct time…
…and then sometimes it simply doesn’t run.
No error. No message. Nothing. The scheduled time just passes.
When it does run, it regularly forgets steps, does things incorrectly, stops halfway through, or produces something I have to manually check and redo anyway.
At some point I thought: okay, maybe one bot shouldn’t be doing all of this.
It’s configured as my Chief of Staff, so I literally convinced it to create a small team just to handle this ONE tiny evening-report workflow.
It created the agents. Created a group chat. Delegated the work. They actually worked on the report together.
And then my “Chief of Staff” simply forgot to collect the results and send me the final report. 😭
I had to message it “how’s it going?” before it remembered that it was supposed to deliver something to me.
I’m genuinely trying to understand what I’m doing wrong here.
Because I really WANT Grok Bot to work. The concept is almost exactly what I want from an agent system.
But after paying $300 for access, recording detailed instructions, building routines, giving it memory/context, and even creating a dedicated team around one simple recurring task… having to constantly remind, supervise, check and redo its work is honestly pretty disappointing.
Am I using Grok Bot fundamentally wrong? Or is this level of reliability currently expected?
Because right now, unfortunately, this has been my worst experience with agents so far.
I’m ~6 months into coding and have shipped 22+ internal tools for real businesses with AI agents. I’m now moving development off production servers onto WSL and learning GitHub, PRs and CI/CD. For a solo builder, what’s the minimum sane setup—and where does “proper engineering” become overkill?
I gave Grok Bot one job today, write me a post for X. The first drafts came out like a translation and then a thesis, so I threw them out. One job still takes a few swings 😂
@thsottiaux I’d love Codex to become one shared brain across all my OpenClaw agents, servers, projects, and tools — with persistent context and the ability to coordinate everything as one system.
Tibo, please don’t do that 🙏 You literally reset the limits just a day ago, and I’ve been trying to save tokens because they were getting consumed way too quickly before due to Sol. I intentionally avoided working on the bigger projects that actually matter to me yesterday so I wouldn’t burn through more tokens than necessary. And now the limits got reset again… when I still had 90% left 😭
I haven't used ChatGPT work much yet, but it's already helped me automate some everyday tasks that don't require as much customization as OpenClaw. Ultimately, these automations are simpler and cheaper❤️🔥❤️🔥❤️🔥 @OpenAI@thsottiaux
@thsottiaux I built my first personal website; brought my first SaaS product to phase 42; and built an automated customer analytics system based on 300+ call recordings from my managers. All this just yesterday. I spent 65% of my weekly quota in one day. 🥲 Help!
I use most models I can get my hands on, Fable included. Different tools for different jobs.
But GPT-5.6 Sol is the one I keep coming back to.
I trust @OpenAI. I like how they build and how seriously they take the work. It feels honest, and it shows.
Thanks, @thsottiaux 🔥
Can someone please explain why @openclaw moved from a stable Pi runtime to a “stable” OpenAI harness that drops every 15 minutes? I keep having to message the agent “continue working” or the whole task just hangs. babysitting…