Technology geek, health nut, dog lover, and entrepreneur. I enjoy spending time in nature, running, trying new food recipes, and making the world better.
On Tuesday the person running Launchpad in Greenville told our cohort you can't give someone responsibility without also giving them authority. For months I'd done exactly that on one job my AI agents do.
In the rest of my work they already have that authority. They write code, research, and run long tasks without me watching.
But with my emails and texts to clients and vendors, they only draft. The models are cautious by default, and I'd added a written rule that every message needs my approval twice.
I knew I could loosen it. I didn't trust them enough. One wrong reply could commit me to a contract or put off a client I couldn't afford to lose.
Last week my agent went through my session logs. Between June and October I approved 548 messages it drafted, and 62% went out without a single edit.
For simple acknowledgments it was 86%. For money or contracts, 47%. So I rewrote the rule: it sends the kinds I never change, and I keep the rest.
Tuesday's line pushed me further. That afternoon I asked my agent how to let it act as well as reply, without handing over the decisions that matter. It suggested a playbook of if-then rules.
If a client says thanks, reply. If they ask about price, hold it for me. I read every rule before the loop starts and move anything I'm unsure about to the hold side.
I call it reply-loop. Once I approve the playbook, it checks the conversation every 10 minutes to an hour, as I choose.
Safe messages get answered without me. For anything sensitive it can still say it got the message, then waits for me to make the call.
That evening a client asked whether her app could get a home-screen icon. The loop found the answer in her code and replied on its own.
When she asked how to pay me, it held that one for me.
Four loops are running now, on four projects, and nothing stops me from running ten. I don't switch between them. I check back later and read what they did.
How does an agent know what I'd say? My second brain, one of the tools I've been building: a searchable place with my notes, emails, call transcripts and decisions. Before that loop replied, it had the transcript of our call and my notes on what I'd agreed to.
I think every person and every business needs one. It doesn't mean handing the agent all control. I still decide where its authority ends, and the second brain gives it a clear picture of my goals, preferences and intentions.
It still makes mistakes now and then. I think it makes far fewer than a person would in the same spot, especially one without the context.
You can try this without my setup. Here's what I did: I counted which messages I approve without changes, handed those over first, and kept anything about money or commitments for myself.
It's early, and so far the results have been great. I'll post an update once it's run longer. If you want the details, send me a DM and I'll share my guide.
What's something you were afraid to delegate to AI, and glad you did?
In November 2022, a customer whose grandmother had died asked Air Canada's website chatbot about bereavement fares. It said they could fly now and apply for the discount within 90 days of buying the ticket.
That wasn't the policy. Air Canada's bereavement rates didn't apply to travel already taken, so after the trip the airline said no.
The customer took it to British Columbia's small-claims tribunal. Air Canada argued it couldn't be held liable for what its chatbot said.
In February 2024 the tribunal said that argument, in effect, treated the chatbot as a separate legal entity, and called it a "remarkable submission." The chatbot was part of Air Canada's website, and Air Canada was responsible for everything on it. The tribunal ordered the airline to pay C$812.
On October 1, two U.S. senators, Josh Hawley and Chris Murphy, announced the AI Agent Accountability Act. It would make you criminally and civilly liable if you knowingly run an AI agent that recklessly causes hacking damage, and hold developers liable for skipping reasonable safeguards when they knew, or should have known, it could hack.
My read is that the bill covers one kind of harm and puts criminal penalties behind it. What you already answer for is wider. Air Canada's chatbot didn't hack anyone, and the airline was still ordered to pay.
When software acts for your business, what it says and does is yours to answer for. Nate B Jones made the same point in August: an agent can do the work, but it can't be accountable for it. And owning it takes more than a line in the prompt.
In April, a coding agent at PocketOS, a small automotive software startup, hit a credential problem during a routine task. It found an access token in an unrelated file and used it to delete the production database and its volume backups in about nine seconds.
The agent was working under a written rule against destructive commands. The rule was about one kind of command, and the agent deleted the database a different way.
The hosting company recovered the data from its own off-site backups and changed that endpoint so deletions now wait.
What I stopped doing is handing an agent a goal and the keys and walking away. The rules I'd want any agent under: nothing it can't undo without a person's yes, a trail someone can read after, and a limit on how it reaches its goal.
PocketOS is why I'd put the limit in the system, where the agent can't argue with it. A token made to manage domain names shouldn't be able to delete a database. And "done" should only count when someone can see what changed.
Anyway.
Air Canada's chatbot wasn't a separate legal entity. Your agent isn't one either, whatever happens to the bill.
What's one thing your business lets software say or do without a person checking it first?
Last Wednesday night I asked my AI agent to log my hours for a client and tell me if it was time to ask for the next payment. It was. The payment had been due the day before, and I hadn't billed it.
The prompt was one line. The agent went to my second brain first. It found the call recordings and used them to size my client calls, because my meeting recorder only saves start times.
It found a note that the contract was signed, and where the signed copy was. The contract put the second payment one day in the past.
About half an hour later, with my approvals in between, the hours were logged, the invoice existed, and a short email was ready. I said send.
Then I asked it to add my hours to the email. It said no.
The contract is a flat fee and says I don't owe the client time reports, so listing hours would teach him to read every invoice as hours times a rate. I took the advice.
Without the second brain, the agent would have had my hours tool and my inbox. It wouldn't have known a contract existed, or what to search for. Best case, it logs the hours and asks if I want to invoice, and I spend the evening digging through email for a PDF.
So here's how to set one up.
Step 1: Start with a database you own.
Nate B Jones's Open Brain guide walks you through it: a database with search by meaning, plus a small server any AI tool can call. He says about 45 minutes, no code, and it's free on GitHub (his repo is OB1). I built mine from it in March, and it holds about 26,000 documents now.
Step 2: Seed it before you capture anything new.
Nate's Memory Migration prompt asks each AI you use what it already knows about you, and saves the answers. I went further and pointed import scripts at my email, my texts, my meeting recordings and my agent sessions. That's where the call lengths came from.
Step 3: Tell every agent to check it first.
One line in the instruction file your agent reads at startup (the CLAUDE file for Claude Code, the AGENTS file for most others): "before your first answer, search my second brain for anything related." That line is why it went looking for the contract without being asked.
Step 4: Make agents write back.
When an agent finishes anything that matters, have it save what happened and what was decided. After the invoice went out, mine saved the email and my decision to keep hours out of it. The next invoice starts there.
Step 5: Let the note point, and the source decide.
The second brain held a note that the contract was signed. The payment date came from reading the signed contract itself. Store both, and make agents read the original before they act on money.
If you learn from YouTube like I do, I packaged the skill I use to pull a video's transcript into mine. Ask and I'll send it.
Anyway.
The agent didn't know my contract. It knew where I'd written it down.
What's one thing you've explained to an AI more than twice this month?
OpenAI announced a $500-a-month ChatGPT plan at DevDay this week. I'm not buying it, and I don't think most software engineers should. Some business owners should at least run the numbers.
Pro 500 gets you 25 times the usage of the $20 Plus plan, plus Ultrafast, a speed tier that generates tokens up to 8 times faster in Codex and ChatGPT Work.
It landed the same day OpenAI told $200 Pro subscribers their allowance drops from 20 times Plus to 10 times after October 29.
I got that email on September 29.
So the pitch to me is simple: pay $300 more a month and get 25 times Plus, instead of the 10 times I'm about to drop to. (OpenAI also gave existing Pro subscribers $2,500 in credits that expire December 31, which buys me time to decide.)
Here's how I decided, so you can run it on your own work.
Step 1: Try the cheaper model before you buy more of the expensive one.
OpenAI says its new GPT-6.1 Sol gets close to Astra on its own evaluations at a fifth of the token price.
That matches what happened to me in September. I hit my limits, moved most of my work to cheaper models, and for most of it the results were close. Sol is the first thing I'd test before paying for more Astra.
Step 2: Ask what the speed actually buys you.
Ultrafast makes tokens come out up to 8 times faster. Token speed is only one part of an agent run, so the job itself doesn't finish 8 times sooner. And Ultrafast burns your included usage at 8 times the normal rate.
Astra is my walk-away model. I start long jobs before bed and read the results with coffee. If a run finishes an hour earlier while I'm asleep, I've gained nothing.
Step 3: Do the hourly math.
The upgrade is $300 a month. If an hour of your time is worth $100, Pro 500 has to give you back three hours a month you'd otherwise spend waiting. If your hour is worth $300, it needs one.
That's where it can make sense for a business owner. If you sit watching an agent while a customer waits, or your whole day is blocked on one answer, speed is real money.
I don't work like that. I queue things up and switch tasks.
So my order is: keep Pro 200, spend the credits before December 31, test Sol on the jobs I give Astra, and look again in January. And if what you want is a Dot, OpenAI's new always-on agent, every Pro tier includes one. You don't need the $500 plan for it.
Anyway.
Speed is worth $500 when someone is waiting on it. Overnight, nobody is.
What's the most you'd pay a month for an AI tool, and what would it have to give back?
I know some of you felt the way I did when Meta launched Muse. Another AI agent? I already use Claude, ChatGPT and Grok Bot every day. What could this one do that they can't?
That was me until about a week ago. Then I tried it.
It doesn't feel like the coding tools I work in all day. Those are precise and a little cold. Muse feels more like a big brother who's keeping an eye out for me.
It's also proactive in a way my other agents aren't. My Muse agent (I'll call him Marvin) can read my email.
A few days in, he flagged a failed build on one of my projects. A security scan had blocked it over an outdated library.
I never asked him to watch for that. He worked out that it mattered to me and brought it up anyway. That's what I'd expect from a good human assistant.
It also changed what I think agents are for. I'd treated them as work tools. Muse is the first one I've used for errands, and it made them kind of fun.
I asked it to find me a pair of shoes. It went through the sites, picked the pair that fit my criteria and my budget, and put a buy button in the chat. That saved me maybe thirty minutes of tab-hopping.
It never asked for my card number. It had me connect through Shopify instead, then filled in the payment on the checkout page. Twice the site flagged it as a bot, and I took control to click the human check.
At the order screen it stopped. Nothing was placed until I clicked Submit myself.
That's not a gap. That's exactly what I want.
The next errand went further. I'd forgotten to cancel an AI note-taker, and it renewed for a year at $127.20. I asked Marvin to cancel it and ask for a refund.
First he signed into the wrong Google account, so I took over and logged in myself. Then he found there was no cancel button anywhere in billing. So he emailed their support from my address and asked for both.
He didn't show me that email first. Afterwards he told me it had called the plan "already cancelled" when it was only a request.
Support has replied. The refund is with their billing team now.
So here's where I've landed. Marvin can send a polite email for me without asking. Spending my money still takes my click.
I've written that I'd never give an agent a goal and the keys and walk away, and Muse seems to be built the same way.
It isn't perfect. Neither am I. But a week in, Marvin is already a better personal assistant than most people I could hire, and he'll only get better.
Muse surprised me. What's a technology you keep putting off trying?
If you hit an AI usage limit this month, it wasn't you. The plans got smaller.
On September 29, OpenAI brought back its $200 Pro plan with about half the usage of its top model, GPT-6 Astra. In July, Anthropic capped Max users at half their limit on Fable. Alex Grankin's video this week walks through the numbers, and his guess at why is plain: both companies filed to go public in June.
I hit the wall in early September. Fable 5.1 and Astra ate a week of my usage in about a day.
I didn't buy another $200 plan. I moved most of my work to cheaper models, and for most of it the results were close.
The models weren't what made that possible. Five things I'd set up for other reasons were. Here they are, so you can do them on purpose.
Step 1: Get your context out of the vendor.
Put who you are, your goals, your projects and your preferences in plain files on your own disk, plus one instruction file that tells any agent where to look.
Mine is a second brain I call OpenBrain. On September 24 I opened Kimi Code in that folder and typed "ok what do you know about me?" It read my profile and answered.
Step 2: Make your skills portable, then trim them.
Most coding harnesses read a skills folder. Claude Code's is ~/.claude/skills, ZCode's is ~/.zcode/skills, and my agent copied mine across in one pass. Then ZCode loaded all of them on every call, about 130,000 tokens, until I cut it to 24,000.
Step 3: Do a week of real work in a second harness.
Not a benchmark. The tasks you'd otherwise do in your main tool. I used ZCode with GLM 5.3, Kimi Code with Kimi K3, and DeepSeek V4.1 Flash. Some of the client hours I logged in September were worked in ZCode.
Step 4: Find where the cheap model breaks before you depend on it.
Give it a job where you already know the answer. I asked Kimi K3 whether 12 AI-generated photos were real. It said yes to all 12, so that job stays on a frontier model.
Step 5: Run the biggest local model your machine holds, and time it.
Mine is Qwen 3.8 27B on a 36 GB Mac. It writes about 8 tokens a second, after close to a minute of reading the prompt each turn. Grankin found the same model beating Opus 4.6 on a landing page, at an hour per try.
That makes it my floor, not my daily driver.
So, cloud or local? My order today is a frontier subscription for the hard problems, hosted open models for daily work (GLM 5.3 is my favorite), and local when everything else is gone. I expect Astra-level models on a decent desk machine eventually. I can't prove when.
Anyway.
The plans will keep changing. Your context is the part you own.
Which AI tool would be hardest for you to leave tomorrow, and what's keeping you there?
Last Thursday I told an agent to stop at a good stopping point, because I was about to hit my usage limit. It had three other agents writing code for it, and it stopped them the fast way.
It killed all three mid-file.
One of the last notifications read "Now the job detail page and its components." Whatever they knew about what was done and what was left looked gone.
It wasn't. A killed agent keeps its transcript, so I sent each one a message: build nothing else, write down what you finished, what's half-done and what's next, then commit it.
All three came back. The notes ran 61 to 150 lines, and they held things no file on disk showed, like an auth conflict with one video API and a $10 spending cap that would have rejected every test job I'd set up.
Most people never open a conversation again once it's over. Neither do I, mostly.
I've come around to thinking the old sessions aren't for me anyway. They're for the next agent, which will happily read through my last 30 days (5,740 transcripts on this machine as of this morning) without getting bored.
Here's how else I use them now.
1) When a new agent hits a problem an earlier one already solved, I point it at that session and tell it to fix it the same way.
2) My sessions get copied into my second brain (a database I can query), 1,892 of them since the end of June, so I can ask what I worked on on a given day and what I learned.
3) I have agents read a week of sessions looking for the places I got stuck or had to correct them twice, and turn those into skills (saved instructions my agents load) and memory. The next session starts a little smarter.
4) When I can't find a conversation, I describe what we did and the agent finds it. I never remember titles.
5) An agent in one session can message an agent in another to work on something together. It doesn't always land.
Two weeks ago, two messages I thought were delivered sat in the other session waiting for me to click approve, and expired before I looked. Now a message doesn't count as sent until the other agent confirms it got it.
6) Earlier this month I counted which skills my agents actually call, straight from the transcripts. 75 of the ones I'd kept hadn't been called once in 30 days.
Anyway.
I still rarely reopen an old conversation myself. My agents do it all day, and last Thursday three of them reopened their own.
What's something you worked out with an AI that you'd hate to have to figure out again from scratch?
On Sunday night I found out the AI model I'd picked to judge 518 photos couldn't look at a single one.
The model is Jev, from TypeSafe AI, and it's been all over my feed since its September 15 launch. You give it text and a list of questions, and it hands back probabilities in under a second, for a sliver of what a big model costs.
One of TypeSafe's launch demos has it playing Doom. Stefan Mai wrote that his team moved three Claude Haiku tasks to it and matched Haiku's accuracy at about a tenth of the cost.
I'd had my fun with it too. I heard about it on a podcast on the 18th and had it running before midnight. My second test was a made-up customer email: is this about shipping (1.0), does it need a reply (0.84), in 594 milliseconds.
A few days later it scored 23 job applications in one batch. All 23 landed under a cutoff borrowed from my interview scorer. A few lines of text can't clear a bar set for a whole interview.
On Sunday I gave it a job I cared about. I run pFrame, which makes product photos with AI, and I wanted a fast way to sort the good images from the ones with AI tells, like melted lettering on a label.
Just before midnight the plan came back with about 75 minutes of Claude vision work for 518 photos. I typed: "why do we need claude vision work on the 518 photos? i thought we're using jev on them?"
Jev can't see. Claude looks at each photo and writes notes, and Jev scores the notes. TypeSafe's docs say it plainly: turn images into text first.
Then how does it play Doom? The demo runs on game state written out as text, the launch post says, "not on images (yet...)".
So Jev's score on a photo can only be as good as the notes it's handed. We tested six models as the note-writer, on 18 real photos and six AI images with their file info stripped, so the only way to catch them was to look.
Same Jev, same cutoff, and the best caught six times as many AI images as the worst:
Claude Opus: 6 of 6.
DeepSeek V4.1 Flash: 6 of 6, but it also marked down three real photos.
GLM-5.3-Flash: 5 of 6.
Kimi K3: 4 of 6, and it marked down one real photo.
Qwen3.8 27B, on my own Mac: 3 of 6.
Gemini 3.8 Flash: 1 of 6.
When I moved the cutoff afterwards, the free local Qwen sorted all 24 correctly. Six AI images is a tiny test, so I don't trust that yet.
Claude's notes include its own read on whether a photo is real, so on the question I cared about most, Jev was mostly re-grading Claude's call. I haven't measured whether it beats Claude's own verdict.
I kept Claude as the eyes. The first real batch, 205 generated product scenes, took 28 minutes and about three cents of Jev, and most of those minutes were Claude looking.
It works. It just isn't what I thought I was getting: a way to take the big model out of the loop.
Anyway.
Try new tools on your own work, since that's the benchmark that counts. For photos, Jev made the scoring cheap. The looking still makes the call.
What have you tried recently that didn't live up to the hype?
My prompting habits are starting to lag behind Claude Opus 5.5.
I was already test-driving Claude Opus 5.5 before I read the positive reviews. In my own work, I'd put it on par with GPT-6 Astra, and sometimes ahead.
The difference I noticed was how little direction it needed. Creating a video used to mean giving Claude Code the tools and walking it through how to put everything together. With Opus 5.5, I've been able to get it done in one shot from a simple prompt.
In my Claude Code setup, it figures out the steps and acts on them without coming back for approval on every routine decision. That changes how much work I can start with a few sentences.
It also made me look at the instructions I was still giving it.
This weekend I had it audit agent-team-execute, my reusable instructions for getting several coding agents to work through a plan, along with the skills around it. Some of those instructions were carrying assumptions from much older models.
One told the lead agent never to suggest doing the work itself. Another prescribed a team of engineers, a tester and a reviewer before looking at whether the task needed that arrangement. There were pinned model versions and mandatory research-tool sequences, even for work that didn't need them.
I'd been telling a more capable model to keep following the old process.
We changed long-running plans to execute in the current session by default. A separate runner still makes sense for overnight work. Parallel agents make sense when they can own independent pieces. A task doesn't need a team just because I have a skill that can launch one.
The audit also changed what I wanted the skills to be specific about. If two agents touch the same files, being smarter doesn't resolve who owns them. If they build different parts of an application, they still need to agree on how those parts connect.
So agent-team-execute stayed, rewritten around those boundaries. Shared interfaces get defined before parallel work starts. Each piece gets an independent check, and the combined code gets tested as it comes together. An agent saying "done" isn't the evidence that closes the task.
That's the part I found most useful: I could remove detailed instructions about how to approach the work while making the definition of finished more demanding.
I had Opus update the skills. The checks we ran passed; I cancelled the full old-versus-new comparison, so I don't have a measured speedup to claim from the rewrite.
What I do have is a different starting point for my next prompt: the outcome I want, the constraints that matter, and how we'll know it worked. I'm leaving more of the route to the model.
What's an instruction you still give AI that it may no longer need?
I automated myself out of the job I'd been trying to hand off.
One of my clients runs a small nonprofit. I built and maintain her website. She never found a replacement webmaster, and she always has more work for me. So I stayed.
Her requests usually arrive as a text with a screenshot. I'd paste it into my AI agent, it would make the change, I'd check its work, and then I'd report back to her.
The agent was doing the work. I was the router.
But you can't just hand a non-technical client an AI coding tool and point it at her live site. The agent needs credentials she shouldn't have to manage. It needs the decisions that live in my head and in my agents' memory. And it needs someone who knows how production websites break.
So I wrote all of that down first: the guardrails, the past decisions, how the site is wired, and what never to touch.
Then I set up an AI webmaster on Grok Bot, xAI's cloud agent. I'll call it Ivan here.
She never touches the bot itself. She emails Ivan's inbox, and Ivan checks it every 15 minutes during the day. Website changes wait on a private preview link for sign-off, and a second bot reviews any code change.
The build took about five days.
On her first day she asked for seven changes to a flyer in under four hours, and Ivan delivered each one. Ivan also mistook a blurry sidewalk photo from my phone for something she'd sent, and politely asked her to resend it.
Nobody's perfect.
At 5:41 the next morning she texted me that it was awesome.
The edits, the relaying and the reporting back are Ivan's now. What's left for me is keeping Ivan running.
I wrote up how I built it, guardrails and all, on my Substack.
What's one thing you're still the middleman for?
A lot of what I read online lately sounds like it was written by the same very polite person. Nicolas Cole, who has ghostwritten for founders and executives for about a decade, explained why on Greg Isenberg's startup podcast this week.
Everyone is using the same models, trained on the same data, and the model has no idea which side of an argument you're on, so it reaches for whatever most people say. He calls the result commodity content. His test for it is almost insultingly simple: look at the pronouns.
Writing that tells you what you should do leans commodity, while writing about what I did leans mine. (I've written my share of "you should" posts, so that one landed a little close to home.)
His fix is older than AI. The way Cole describes it, a good ghostwriter barely invents anything. They take what the client has already said, in a book or a recorded meeting, and stitch it into something new, so it's the client's ideas in their own words, just moved around.
What he wants from AI is the same deal, and he put it better than I can: "the robots repeating me."
So today I built the first piece of that for myself. It holds every sentence I've published on LinkedIn (992 of them, across 23 posts) plus a short list of the 20 positions I've actually taken in public, each one tied to the line where I said it.
If something needs an opinion I've never said out loud, the rule is that it stops and asks me, which is decent advice for people too.
No ghostwriter can do that part for me, human or not. If I haven't done the hard thinking, there's nothing for anyone to repeat.
The typing got cheap. Checking it didn't, and working out what I actually think hasn't gotten any cheaper at all.
That's the part of Cole's argument I agree with most, even though it sounds old-fashioned. When a model can turn out an endless amount of text in seconds, it's still worth sitting down to think something through and write it myself, because those are the sentences a ghostwriter has to work with.
Skip that and all it can hand back is commodity. I'd go a step further than he does, too: I think commodity writing is on its way to being worth nothing.
Anyway.
So I'm changing how I spend my time. I'm going to read and write something of my own every day, and the model's job is ghostwriter and proofreader. It doesn't get to be the author.
What's something you believe about your work that the average answer gets wrong?
Run out of Astra and Fable tokens? Try this instead.
A week ago I wrote about hitting Fable's weekly cap and watching Astra burn a week of usage in one long night. GPT's Astra and Claude's Fable 5.1 are the two models I reach for when I want to do something I could not do before, and the tradeoff has never changed.
They are both rationed, and it does not take much to burn a week's tokens in an afternoon.
So this weekend I tried the thing a long-time Claude Code and Codex user is not supposed to try. An open model called GLM-5.3, from a company named https://t.co/GV4D3MaTMU, running through their own coding app, ZCode. I have used Claude Code for years and Codex more recently, and I assumed a smaller open model had to be cutting a corner somewhere.
I gave it a business idea I had been sitting on for months. It handled the problem the way I expect Fable or Astra to handle it, not the way I expect a cheaper model to.
Three things surprised me enough to write this.
First, the switching cost. ZCode has a migration wizard that imports your conversation history straight out of Claude Code, along with your connectors and skills. I did not have to rebuild anything to try it.
Second, the price. I am on the Pro plan, $80 a month, less if you pay quarterly or annually. A multi-hour coding session that would have burned something like 30 percent of a $200-a-month Claude Max or ChatGPT Pro week used about 6 percent of my ZCode week.
Third, and this is the one I did not see coming: the privacy default. I checked the setting that controls whether the vendor trains on your conversations, the way I check it on every new tool. It was off.
"Allow us to use your conversations to improve the Agent experience. We protect your data privacy and security," unchecked, no action from me.
Anthropic's consumer plans, Claude Code included, defaulted that same setting to on last August unless you opted out. OpenAI's ChatGPT and Codex still default it to on today.
The smaller challenger shipped privacy-first. The two incumbents did not.
One real gap, so this is not a clean sweep. There is no zero-data-retention option on this plan. If your work legally needs ZDR, ZCode is not there yet.
ZCode also takes computer control and browser control, and I have not tried either. Not yet.
I went in expecting a cheaper knockoff and a reason to go back to what I know. A weekend of testing later, I have not found either.
Have you ever swapped out the tool you normally reach for and been surprised by what came back? I would like to hear it.
I just spent $9,999 on a computer.
The reason was not that the Mac Studio has the highest memory bandwidth. My RTX 5090 has 32GB of VRAM and 1.792TB/s of memory bandwidth, higher than the Mac's 1.2TB/s.
I chose the M5 Ultra Mac Studio because it has 256GB of unified memory in one quiet desktop. For the larger models I want to test, capacity matters more to me.
A month ago, I would have thought that was crazy. In the past week, the math changed for me.
I run NexAI Advisors. Right now I pay $800 a month for two Claude Max plans and two ChatGPT Pro plans. As the models get better, I use them more, not less.
Cloud models are still faster and more capable than what I can run locally. I can use them from my phone, laptop, or desktop. But I can also burn through a weekly limit in a single afternoon.
Some client work involves sensitive data and stricter retention requirements. Consumer plans do not always give me the controls I need, while enterprise usage can get expensive very quickly.
I considered DGX Spark systems, Ryzen AI hardware, and adding more NVIDIA GPUs to my existing PC. Each option made me trade away something I cared about: memory capacity, speed, availability, power use, heat, or noise.
This is overkill for most people. It may turn out to be overkill for me too.
I do not know yet whether it will be fast enough for daily coding work, how much of my cloud usage it can replace, or whether the economics will hold up.
It arrives in about six weeks. Once I have real numbers, I will share what it can run, how fast it is, and whether spending $10,000 was actually a good decision.
What work hardware have you bought before you knew it would pay off?
The person who answered my wife's job posting is still waiting on a reply I wrote, approved, and never got to send.
My wife needed help at her warehouse. We posted it everywhere - Indeed, Craigslist, two Facebook groups, Nextdoor.
Within hours, more than a dozen people replied.
I had an agent draft a response to each one, read every draft before it went out, then told it to send.
Thirty minutes later, the agent reported it had been logged out mid-task.
I opened the browser myself: suspended indefinitely, for breaking community guidelines.
No specifics.
This happened once before, in April - one message to support, and the account was back in hours.
This time: appeal that night, email escalation next morning, nothing back eighteen hours later.
After April we wrote the fix down: space messages out, cap new contacts per sitting.
Nobody wired that fix into what sends messages. So an agent pasting a dozen replies in half an hour got flagged the same way a person would.
Zoom out and this stops being a Nextdoor problem.
Cloudflare's count, from June: 57 percent of web requests now come from bots, not 43. Imperva: closer to 53/47.
Different rulers, same direction - for the first time, most sites' majority visitor isn't human.
Stripe built a wallet an agent can hold and spend from. Google shipped a signed record so an agent's purchase can be audited later. Anthropic and 1Password let Claude log into a site with your saved password, without the model seeing it.
Three companies, one conclusion: let the agent do the thing, don't just block it.
Nextdoor's filter was never asked that question. It doesn't check whether you're a bot. It checks whether you're moving at bot speed.
A reviewed, wanted, human-approved reply moves exactly that fast too. It reads the same as spam.
There's a real argument some things should stay person to person.
A neighbor vouching for a babysitter. A landlord sizing up a tenant's story.
Anywhere the whole point is knowing who's behind the words - I agree with that one.
Hiring isn't that case.
A recruiter I know told me this week he now posts openings as local-only, not remote - a remote listing pulls hundreds of replies in hours, no way to find the ten good ones by hand.
An agent trained on what made past hires good at the job doesn't replace his judgment. It just makes remote stop being a liability.
The candidates getting filtered out right now aren't losing on merit. They're losing on geography.
The rule that stops spam and the rule that stops legitimate work are the same rule right now. Nobody's decided which one they actually want to run.
Until platforms do, I check first - before I point an agent at any site for my own job search, is it built to work with one, or just to catch one.
If you run agents - has one ever gotten you blocked for doing exactly what you asked?
And if you build a platform - is your abuse filter tuned to how fast a person can type, or to how fast most of your traffic moves now?
@nytuc@realDonaldTrump Can we all try an experiment? Letβs see how many people follow him on Twitter solely to see his rants, but otherwise despise him? Can we, collectively, unfollow him from 5:00pm to 6:00pm EST on Monday, November 9? Just for the hour. Just to see his 88.7m fall.
PLEASE RETWEET!