Not a developer. A decade in tech, analytics, strategy and business. Now I use AI in my own projects and post what worked, what it cost, what you can copy.
I gave Jev, an AI model that answers with probabilities, 41 of my real X replies and asked which ones the author would answer. 15 had.
It said no to all 41. That scores 63%, the same as always guessing no.
But its ranking was real. Of the 10 it rated most likely, 5 got an answer. Of its bottom 10, 1 did.
32 seconds, under a tenth of a cent.
As a yes or no judge, useless here. As a sorter, worth a second test.
The receipt that changed my mind wasn't a delivery success. It was zero replies for a week.
I stopped asking "did it send?" and started asking "did it land where they already look?"
If you moved one report from email into WhatsApp or Slack, what broke first, formatting, noise, or permissions?
I built a clean automation that emailed the daily summary.
Status said sent. Job looked done.
Nobody opened it. No replies. No decisions.
The owners were already in WhatsApp and Slack. Email was where the report went to die.
I moved the same output into the chat they already open.
Same facts. Different place. Suddenly people replied.
"Sent" is not the same as "seen where they work."
Where do your "done" automations land, inbox, or the chat they actually live in?
The fail I care about wasn't a wrong answer. It was a run I never meant to buy.
"Smart by default" looks helpful until you check the receipt.
Curious: do you gate the strong model with a typed yes, a money cap, or something else?
I used to let Fable (the expensive model) decide when a task was "hard enough" to use it.
It picked itself too often. Quiet bills.
New rule: Fable does not run until I type yes.
Cheaper models handle drafts, sorting, first passes. If I'm unsure or away, Fable waits.
Slower on purpose. Cheaper. Fewer surprise tokens.
Where do you still let the expensive model decide for itself?
@DaveComeAtMe Hosting. I set a server to stay on, the product never launched, and it ran for months anyway, every invoice $0 so nothing flagged it. The gap was not a crash, it was not knowing it was still running long after I moved on.
@Jasonturcotte Payroll is a good first one because it repeats. Set it to shout when the newest data is older than a day, not only when a run fails. Mine was switched off for 7 days and nothing told me.
@DaveComeAtMe What caught me was the step after the spec. Claude told me a batch of images was ready to go out, I opened them, and all 7 were black rectangles at the wrong size. Nothing counts as done here until someone opens the real output before the customer does.
Before you automate anything, 5 things:
1. Do the job by hand once and write down how many minutes it took.
2. On the same page, the one number this job is supposed to move.
3. Run it once with Claude or ChatGPT and compare the two results line by line.
4. Find out what one run costs you, and write down where you read that figure.
5. Put a date in your calendar to check whether the number moved.
What would you add?
@bmtriet Good use. Worth opening the device list after a reboot before you call it fixed. Mine reported 3 emails sent that never arrived, so I check the thing itself now instead of the report that it worked.
@Abir2 I ran my own posts through 6 rounds of AI fact-checking and got them back accurate, short and dead. What I do now is write the thing from my own figures first and hand it over only for the fact check. The check judges claims, it cannot tell whether anyone learned anything.
@JoshExile82@Muse@typesafeai The way I stopped relying on it choosing to check: nothing gets filed unless the check left a receipt on that exact text. No receipt means it never ran, and the work does not go out. It turned did it check into something I can look at.
@charlynortonn The next thing to check is a site that shows up in Google but not in ChatGPT. My own did that: ChatGPT got a title and a list of links where 6,773 characters of page should have been. Start with the pages that bring in the money.
THE AI PROJECT COST CARD
by @shaharfrish · free, copy it, change it, no sign-up
For a business owner moving one task to AI. Fill it in before you start, and again after 30 days. It exists to stop the two quiet failures: something runs for months that nobody wanted, and something reports done when it produced nothing.
BEFORE YOU START
Write a real answer next to each one. If you cannot answer it, that is the answer.
1. The one task I am moving to AI: ______
2. What that task costs me today: ______ in money each month, and ______ minutes each week
3. What a wrong answer from it costs me: ______
4. Who reads its output before anyone acts on it: ______
5. Done means this exists and I can open it: ______ (a file, a row, a message in the sent folder)
6. The number I expect to move, and the date I will check it: ______ by ______
7. What I will stop doing if it works: ______
8. Who can switch it off without asking anyone: ______ . Where its failures will show up, in a place I already read: ______
LABEL EVERY FIGURE BEFORE YOU WRITE IT DOWN
- Money out. Cash that actually left your account.
- Time. Minutes you actually spent.
- Damage. Something already went wrong and no cash moved.
- Hazard. Nothing has gone wrong yet, and it is set up so that it could.
A hazard is not a cost. Never add one into a total. Two of mine. A template I had staged carried one line that would have sent every request through a third company's server. I caught it before it ran. That is a hazard, not one dollar. A server of mine ran about four months on invoices that all read $0. That is something nobody was watching, not money lost.
ONE FILLED IN: A KEEP
The task: answering my own routine email.
- What it cost me before: ______ in money, and ______ minutes a week.
- What a wrong answer costs: a wrong reply to an important person, which I cannot take back.
- Who reads the output before anyone acts: me, on everything that matters.
- Done means: a reply exists in my sent folder, in my voice.
- What I stopped doing if it works: writing the routine replies myself.
What I did: two passes a day, morning and evening. For the first week it was only allowed to prepare drafts, and I sent every one by hand. It made no mistakes that week, so it earned one step: the lowest-priority mail it now answers on its own. The two tiers above still come to me as a draft and I send them.
The decision: keep it, with the ladder in place. The one thing it does alone is the tier I handed it on purpose, after a week of reading every draft. Its replies are in my sent folder. Nothing important goes out before I have read it. What earned the keep was not the AI getting better, it was giving it one tier at a time, and only after a week with no mistakes.
THE FIVE COSTS NOBODY COUNTS
Five fields. Put a number or a blank in each one, and run the check next to it.
1. The thing nobody is watching. ______ Check: once a month open your provider account and list what is running. Do not read the invoice. A $0 invoice is not the same as nothing running.
2. Work nobody asked for. ______ Check: every automation you own starts because someone asked or a date came, never because it had nothing else to do.
3. Top price for typing. ______ Check: split the work into judgment and typing, and pay top price only for the judgment.
4. A confident wrong answer. ______ Check: any step whose answer the next steps build on, like a date or a customer number, goes to your best AI. Check that answer against the source, not against whether it looks right.
5. The minutes you spend checking. ______ Check: confirm the thing exists, not that the run finished. Open the file, the row, or the sent message.
ONE FILLED IN: A REPAIR
The task. A helper that searches for people worth answering and mails me the list, every two hours from 10:00 to 22:00.
What it cost me before. ______ Blank. Write this one down before the tool goes in. Afterwards it is a guess.
Money out. $0.52 a run. It hit its own ceiling partway through a run, so the last third came back thin.
Time spent checking. ______ Blank until you time it. A feeling is not a figure.
The rework. One evening it returned 2 names. Both were replies buried inside other people's threads, so both were unusable and I skipped them. That night's supply was zero. For a stretch it sent empty reports while I waited, and whole blocks of the day went by with nothing done. Once I waited 3 hours for a report that was never coming, because the schedule had never been switched on at all. Its own mails were not landing at all. A reply inside its thread came back in 48 seconds with that day's three reports.
The error rate I could live with. ______ Set this before day one. Without it, two unusable names out of two is just an evening, not a signal.
The decision: repair. The work was real and the delivery around it was broken. Three fixes beat switching it off. It now brings back only posts that start a conversation, never replies buried in someone else's. A second way to find the work took over when it returned nothing. Its reports now arrive as replies inside one long thread.
THE KILL LIST
Six yes or no questions. Each answer says keep it, repair it or kill it. Run them on day 30, and once a month after that.
1. Can you open the thing it made today, right now? No means repair it.
2. If it had failed every day this week, would you know by Friday, in a place you already read? No means repair it. A daily job of mine failed roughly 5 days out of 6 for a full month. The alert fired correctly every single time and I still found it by accident, a month later, while building something else. An identical banner every morning stops being read inside a week.
3. Can someone inside your business switch it off alone, and do you know what still bills after you do? No means repair it.
4. Does it act alone on anything you did not hand it on purpose, after a trial where a person read every answer? Yes means repair it: take that step back until it has earned it.
5. Do you know what one run costs, and the point where it stops working? No means repair it.
6. If it vanished tonight, would anybody ask for it back? No means kill it.
WHAT GOOD LOOKS LIKE AFTER 30 DAYS
You can list everything you have running, and name one dated thing each of them produced.
When it breaks, you find out from the place you already read, not by accident a month later.
The number you wrote down on day one moved, and you have actually stopped doing the thing it replaced.
@Sabrina_Ramonov For me the change that paid off was not the brand, it was adding a second pass over every draft. I have it check each number against the file it came from. It caught one of mine that said 1 where the file said 15.
@MarsHomestead The moment money moves I want a tick from a person first. Mine posts the proposal to a shared checklist and waits for me to tick the box before it starts. The order stays yours, the machine just does the typing.
@densancar Worth checking what else AI cannot see. On a store I work on, 279 products never showed up for customers at all, and most images had no text description behind them for search. Both were free to fix.
@connorgallic For 3 years my job was one question: did the new thing move the number. Same here: put it on part of the calls and compare bookings and hang-ups, not opinions about what people over 60 accept.