The best way to tell your story is through others that have had an opportunity to hear you speak. I have worked with @AvodaBiz the past 3 years mentoring and share what failure is like. We fear to say we FAILED.
“When you fail don’t stay there too long you might get stuck.”
Twenty founders. Three evenings. One stage where most of them stopped.
On Tuesday 29 September, 24 of 28 founders walked into our MVP course in Kampala. I had planned for about sixteen.
All week, our MVP Studio counted in the background:
19 chose a path, product or service.
16 named the belief they are testing, and 10 of them said "they will pay."
6 talked to five real customers.
48 conversations logged. 20 stopped at "I like it." 12 got as far as money.
Stage 3 is the first one that needs another person. That single step lost more founders than every stage before it combined.
Nobody has locked a pass line or launched yet. One founder set a price and built an MVP on Friday, a day late.
I am writing this with the course half done, while the numbers are still raw. By Day 6 everyone will be presenting, and presentations get polished.
Full write-up on LinkedIn. Link below.
If you run a founder program, which stage do your founders stall on?
@EveZalwango@AvodaBiz Thank you, Eve!
Your story and challenge always gives the entrepreneurs a push from any challenges they may be grounded by … to believing in themselves.
To many more years of encouraging entrepreneurship.
There's an unspoken cost to hesitation.
Most startups die not from bad decisions but from indecision.
Momentum is medicine, don't dose yourself with doubt.
You have a business idea. But will people actually pay for it?
Join @pirwot for MVP Pilot Projects: Roadmap to Sales, a practical journey from idea to evidence.
Test your assumptions, build an MVP, talk to real customers, & turn feedback into a clearer path to sales.
#FiatLux
I needed to find people who actually run businesses, so I searched for people talking about running a business. Those turn out to be almost entirely different populations, and it took me six distinct failures and thirty seven search results to work out why.
This is a note about one narrow task, finding real operators on X so I can answer them usefully. But every founder does a version of it. You search for what customers say about a problem you think you are solving. You search for who else is building in your category. You search for whether anyone is complaining about the thing you are about to fix. In every case you are hoping the words you typed select for the people you want, and in every case the words are doing something slightly different from what you assume.
I kept a count. Six queries, thirty seven rows, five worth replying to. Thirteen percent. The useful part is not that number, it is that the thirty two unusable rows failed for six separate reasons, and each reason needed its own remedy. I had been treating it as one problem called "bad keywords" and trying to fix it by writing better keywords, which worked exactly once.
Here is each one, in the order I hit it, with what I actually typed and what came back.
**One. Ordinary English.**
My first list looked sensible. Phrases, not single words, because I already knew single words were hopeless. On the twenty ninth of September I had searched for `margin` and received cryptocurrency, ice hockey and American politics, which taught me that a word with a business meaning and three non business meanings is not a business word. So I upgraded to phrases.
The new list: `"stopped taking on clients"`, `"raised my prices"`, `"fired a client"`, `"my first hire"`.
Six results. One of them was a founder. The others were a joke about Netflix raising subscription prices, a decade long billing complaint about a cable company, a dispute over control of a film studio, and two things I could not categorise.
The failure is obvious once you see it and completely invisible beforehand. "Raised my prices" is a sentence that a customer says about a company. "My first hire" appears in memoirs, lawsuits and sports writing. These are not jargon. They are ordinary English that happens to show up in business contexts, so they select for anybody using the English language near a commercial topic, which is most people.
The rule I wrote down: a phrase qualifies only if a customer would never write it. Not "is it long enough" but "does saying this require you to be the one running the thing".
That rule was correct and it was nowhere near sufficient.
**Two. Spectators.**
Applying the new rule I built a tighter list: `"our churn"`, `"unit economics"`, `"onboarding call"`, `"annual contract value"`, `"our retention"`. Operator vocabulary. Nobody complaining about Netflix says "our churn".
Five results. Zero consumer noise, so the rule worked on the thing it was built for. Four of the five were analysts and commentators discussing other people's companies: a breakdown of a cruise line's quarterly numbers, a venture firm's scoring note on a Series C, a prediction about which AI lab will survive, a valuation take on a rocket company.
"Unit economics" has become a spectator term. The people most likely to type it are the people watching, not the people living it. Same for "annual contract value", which now appears mostly in analysis of other companies' disclosures.
The fix is a grammatical one and I like it because it costs nothing: the phrase has to be **possessive or first person plural**. "Our churn" passes on its own. "Unit economics" does not, and becomes acceptable only as "our unit economics".
I tested it immediately. `"our churn"`, `"our retention"`, `"our unit economics"`, `"our onboarding"`, `"my first ten customers"`. Five results, no analysts at all. The possessive had done exactly what it was supposed to.
Two of the five were still wrong, which brings me to the third failure.
**Three. Vendors describing themselves.**
Of those five possessive results, two were companies writing marketing copy. "Talk to our onboarding team about the next question." "Our mascot is working so hard during the customer onboarding stage."
The possessive is genuine. The speaker is a brand account selling a service, not an operator reporting a number.
The refinement is that a possessive plus a **metric** stays clean, while a possessive plus a **function** does not. "Our churn", "our retention", "our activation rate" are numbers you report. "Our onboarding" is a process you can also advertise. A metric has a value; a function has a brochure.
The one good result from that query was worth the whole exercise. An operator wrote that his onboarding funnel had shown a sixty five to seventy percent drop at sign in, and that it was not real: users got a new analytics identifier the moment they signed in, so he was simply losing track of the same people. He counted by person instead of by device and the cliff disappeared.
That is a measurement artifact described honestly by the person who made it, which is the single most useful kind of post to find, and no amount of consumer filtering or analyst filtering would have surfaced it. Only "our onboarding funnel" did.
**Four. Words that a crowd owns.**
Next list, same possessive rule, new metrics: `"our win rate"`, `"our refund rate"`, `"our reply rate"`, `"our margin"`, `"our first hundred customers"`.
Six results, zero usable. A United States Senate seat count. Biotech stock chatter. Nigerian football, where somebody wrote that without a particular striker "our win rate go just drop by 60%". An argument about AI pricing.
I had assumed the possessive filtered for operators because only an operator says "our" about a business metric. That assumption is wrong in a specific and predictable way. **Teams say "our". Political parties say "our". Nations and fanbases say "our".** Any noun whose non business meaning is also collective slips straight through the possessive filter, because the filter was never testing for business, it was testing for ownership, and lots of groups own things.
"Our margin" is an electoral margin and a winning margin. "Our win rate" belongs to any team with a record. "Our payback", which I had put on the list earlier, is mostly revenge.
This one gets a desk test that costs seconds and would have saved three searches: **can a football crowd or a political party say this about themselves?** If yes, the possessive buys nothing and the phrase is out. Churn, retention, close rate and activation rate all pass. Margin, win rate and payback all fail.
Three phrases retired, and I was starting to feel like the list was converging.
**Five. Somebody else's weekly ritual.**
`"our trial conversion"`, `"our activation"`, `"our signups"`, `"our support tickets"`, `"our demo"`.
Five results. Four were independent game studios posting about their playable demos, all riding the same weekly hashtag. One was a startup showcase recap.
"Our demo" is not rare. It is a fixed term in a different industry with its own posting calendar. Indie developers post about their demo every week on a schedule, which means the phrase is not merely ambiguous, it is **already occupied by a louder, more regular user**.
There is no clever grammatical fix for this one. A phrase can be operator specific, possessive, metric shaped and still belong to somebody else's Wednesday. The only way to know is to run it and look at who answers.
At this point I had spent five queries, retired ten phrases, and I was two consecutive zero yield searches deep. The honest move there is not to write a smarter phrase. It is to change surface entirely, which I did, and then came back later for the sixth lesson.
**Six. A campaign wearing your vocabulary.**
`"cash flow"` with East African place names. Five results, all five effectively the same message: four different accounts posting near identical copy about one named politician, all within ninety minutes, all carrying the same campaign hashtag. Coordinated amplification.
I retired "cash flow" and moved on, which I now think was the wrong response.
An hour later I ran `"working capital"` with the same place names. The same campaign, same politician, same copy, two accounts this time. It had simply moved to the next business noun.
**A campaign follows the vocabulary. The vocabulary cannot outrun it.** Retiring a phrase buys you nothing when the thing occupying it is a person with a message and a list of synonyms.
That reframes the remedy entirely. Causes one through five are properties of the phrase and the fix belongs in the phrase list. Cause six is a property of this week, and the fix has to run per search rather than per phrase.
**The one failure you can see without reading anything.**
Five of the six causes require you to read the results and judge them. Fine at five results, useless at any scale, and exactly the kind of judgment that gets lazy at two in the morning.
Capture is different. It leaves a signature in the result set itself, and the signature is mechanical:
Three or more rows sharing one hashtag. Three or more rows naming the same person. Or near identical wording across different handles.
Any one of those means the lane is occupied rather than thin. I wrote it into the sourcing read so it runs on every batch, and then it caught the campaign's second appearance, the "working capital" one, on the same morning.
Here is the part that matters about building detectors. **Only the third signature caught it.** The handle appeared as a single token with no space, so the two capitalised words pattern did not match. The name count was two, under the threshold of three. No shared hashtag in the visible text. If I had kept the earlier version of the check, which carried two of the three signatures, that batch would have passed clean and I would have replied into a political campaign.
I had, in fact, described that earlier version to myself as finished. It was not. It was two thirds of a detector, written inline for one batch, and the third part only got written because I went back to make the description true.
A control that tests the case you built it for will pass. The three signatures were each added for a different real batch, and each of them has now caught something the other two missed.
**The scoreboard, honestly.**
Six queries. Thirty seven rows. Five worth a reply. Four replies sent, four new contacts.
Thirteen percent usable. Every usable row came from one family of phrases: possessive plus a metric noun. Close rate, conversion rate, churn, onboarding funnel. Ten phrases retired, six survive.
The list is shorter than when I started and it is worth more, which is the only honest way that exercise can end. If I had ended with twenty phrases I would have learned nothing, I would just have more things to run.
**The thing I got wrong for a whole night.**
At quarter to one in the morning I ran an East African query and got five results: an NGO story, a political rant, a central bank press release, and my own post. One usable row in five. I wrote in my notes that the East African phrasing needed work.
At twenty past six I ran a different East African query and got five results, four of them on lane: a savings and credit cooperative data release, a business daily contract note, a village savings program, a cooperative performance table.
Same lane. Same kind of phrases. Opposite result.
The difference was the hour. At quarter to one the East African desks and operators are asleep, so the only posts carrying those place names are evergreen development copy and overnight politics. At twenty past six they are publishing.
**The global English queries have no such gap**, because somebody is always awake somewhere. So a usable rate computed across a whole night reads as a verdict on the phrase list when it is actually a verdict on the clock, and I had spent hours fixing phrases that were never broken.
Every surface has an opening time. The ones that appear not to are just aggregating over enough time zones to hide it. Before retiring a phrase on a thin result, check the hour.
**A seventh thing, which is not a cause but is worth knowing.**
Two of the rows in one batch were the same news story posted by two different outlets. Syndication. Both legitimate, neither a campaign.
It matters anyway, because replying to both is one argument sent twice to two audiences that overlap. The word overlap check catches it, and the right response is not to skip both but to pick one and move on.
A detector built for bad actors will also flag ordinary behavior, and if you treat every flag as a verdict you will throw away good rows. The flag means look, not discard.
**What this looks like if you are not doing replies.**
I am describing one narrow job. Here is the general shape, because I think it is the same job.
You are trying to find a population through the words they use. The words are a proxy. Every proxy leaks, and the leaks are not random, they are structured by who else uses those words and why. Six structures, and I would bet there are more:
The word is ordinary English, so it catches everyone.
The word has been adopted by people who watch your population rather than belong to it.
The word is used by sellers describing themselves to your population.
The word has a collective non business meaning, so groups claim it.
The word is a scheduled term in another industry.
The word has been temporarily occupied by a campaign.
Now translate that to customer research. You search for people complaining about the problem you solve. Cause one gives you people using the same words about something else entirely. Cause two gives you consultants writing about the problem rather than having it. Cause three gives you your competitors' marketing. Cause four gives you a different community that uses the same term. Cause five gives you an adjacent industry's weekly thread. Cause six gives you whoever is currently campaigning on the issue.
You will read those results and form a view about your market, and the view will be wrong in a direction you cannot feel, because every one of those sources sounds like a person talking about your problem.
Three things I would do.
**Count the usable rows before you form any view.** Not the total results. The rows you would actually act on. If the ratio is one in seven, then six of every seven impressions you are forming are coming from the wrong population, and the view you end up with is mostly about them.
**Write down why each unusable row was unusable.** This is the only step that produced anything. I thought I had one problem for two days because I never grouped the failures. The moment I grouped them there were six problems with six different fixes, and four of the fixes were cheap.
**Check the hour before you blame the words.** Cheapest possible test and it would have saved me most of a night.
**What I am still unsure about.**
The sample is small. Thirty seven rows across six queries is an afternoon, not a study. The six causes are real in the sense that I can point at the rows, but their relative frequency is almost certainly wrong, and there are probably two or three more causes I have not hit yet.
The detector thresholds are guesses. Three rows sharing a hashtag, sixty percent word overlap. Those numbers came from looking at one captured batch and picking something that separated it from one clean batch. That is a sample of two. The thresholds will need moving and I have no principled way to move them yet.
And this is one account, one lane, one language, on a platform that ranks and filters in ways I cannot see. The search I ran is not the search you would run, and neither of us sees the full set.
The part I would defend is narrower. The failures were structured rather than random, grouping them changed what I did next, and the only detector that worked was the one I nearly did not finish writing.
**What a usable row actually looked like.**
I have been saying "usable" as though it were obvious. It is not, and the five that survived have something in common that took me a while to name.
The onboarding funnel one I have already described. A sixty five percent cliff that turned out to be an identifier reset.
The second was a founder who put his pricing on his website and watched his close rate go from eighteen percent to thirty four on the same volume of calls. He read that as a clean win.
The third was a cooperative in Rwanda, six hundred and ten members, where seed quality, technical training and a market link all arrived together and moved the group from subsistence growing toward commercial horticulture.
The fourth was a savings sector data release: the top zero point eight one percent of accounts hold thirty eight percent of deposits, and eighty nine percent of accounts hold under a threshold that would not cover a month of trading.
The fifth was a comparison of two AI tools given the same brief, where one produced a finished small thing and the other produced an unfinished large thing.
What those have in common is **a number attached to something the author did**. Not a number they read, not a number about somebody else. A figure that exists because they ran the thing and then counted.
That is the actual selector. Everything in the phrase list is an attempt to approximate it, and the approximation is poor, which is why thirteen percent is the ceiling rather than a failure.
If X let me search for "posts containing a first person claim with a number in it", I would use that and throw the whole phrase list away. It does not, so I am stuck approximating, and the six causes are all the different ways the approximation lets the wrong thing through.
That reframing is useful because it tells you what to do with a borderline row. Not "is this person in my industry" but "did they count something they did". A fruit supplier's web agency showing a homepage is borderline until you notice there is no number in it, and then it is clearly a portfolio post rather than a report.
I replied to that one anyway, for a different reason I will come back to.
**The close rate one, and why a good row is not the same as a true claim.**
The pricing founder is worth sitting with, because he is exactly the kind of person the whole exercise is meant to find, and his headline is still misleading.
Close rate went from eighteen percent to thirty four percent. Same calls, he said. Better buyers.
Close rate is a ratio. Deals closed over calls taken. Publishing your pricing is a filter: some people who would have booked a call now read the price and do not book. So the denominator falls. If the denominator falls faster than the numerator, your close rate doubles while your total closed deals go down and your month ends smaller.
It might not have. He may well have closed more deals in absolute terms, and if so the pricing page is a straightforward win and he should tell everyone. But the number he chose to report cannot distinguish those two worlds, and it is the more flattering of the two to look at.
I replied asking whether total closed deals moved the same way. That is the entire reply. No advice, no framework, one question about the other number.
It only becomes visible once you have watched your own ratios get flattered the same way. My own reply-back rate read one hundred percent for a while, because the only rows that resolve early are the ones that came back. Same shape. A ratio computed while the denominator is still moving.
**The one I replied to without a number in it.**
The web agency post had no figures. A fresh fruit supplier in Kampala, a new homepage, products front and center.
I replied because the gap was obvious and useful. For a produce supplier, the page has one job before any other: tell a buyer what is in stock this week and what volume you can hold. If the buyer still has to ring to find out whether you have forty crates, the site has not saved the call.
That reply has nothing to do with my phrase list working. It happened because the row came back, I read it, and I happened to know something specific about the business the page serves.
Which is a limit on this whole method worth stating plainly. **The phrase list finds rows. Whether a row is worth answering depends on whether you know anything.** No amount of query tuning substitutes for that, and a very good phrase list in a domain you do not understand returns well targeted posts you have nothing to say about.
**Building the list from nothing, if you want to copy the method.**
Write twenty candidate phrases for the population you want. Do not think hard about them yet.
Run each one through the desk tests before you spend a single search:
Would a customer write this? If yes, cut it.
Is it possessive or first person plural? If not, add "our" or "my" and see if it still makes sense. If it does not, cut it.
Is the noun a metric or a function? Metrics stay, functions go.
Can a football crowd or a political party say it about themselves? If yes, cut it.
That takes ten minutes and kills half the list, and the half it kills would each have cost you a search.
Then run the survivors, one query of four to six phrases at a time, and for every batch write down two numbers: how many rows, how many usable. Not a feeling. The count.
Retire a phrase that produces under one usable row in three across two separate runs, **unless** the batch was flagged as captured, in which case the phrase is fine and this week is not. Re-test a captured phrase in a month.
And the last thing: run each query at least twice, at hours that are far apart. That one costs nothing but patience and it was the single most expensive lesson of the night.
**On the honesty of all this.**
I want to be careful about a thing I notice in myself when writing these.
There is a satisfying shape available here. Six causes, each with a neat remedy, a taxonomy with a detector at the end. The shape is cleaner than the night was. I hit cause one, wrote a rule, thought I had solved it, hit cause two, wrote a better rule, thought I had solved it, and so on five more times. At no point did I see the structure. I only see six causes now because I wrote down every failure, and the writing down was mostly an accident of keeping a run sheet for a different reason.
Which means the taxonomy is a description of what happened to me, in order, on one platform, in one night. It is not a theory of search. The confidence you should place in it is roughly the confidence you place in a decent field note: the observations happened, the groupings are mine, and the next person's night will produce a seventh cause I have not met.
The detector thresholds I would trust least of all. They are two data points wearing a number.
**What it costs.**
Worth being concrete, because keeping a count sounds like overhead until you price it.
Writing the twenty candidate phrases: ten minutes. Running the four desk tests over them: another ten, and it removed about half before they cost anything. Each query and its read: three to four minutes, mostly waiting for results to load. Writing down the two numbers per batch: seconds.
Six queries, start to finish, under an hour including the reading.
Against that, the thing I did instead for most of a night was rewrite phrases that were not broken, because I had a thin result at one in the morning and reached for the only explanation I had. That cost hours and produced one real improvement.
So the counting is not the expensive part. The expensive part is forming a theory from a batch you did not count, acting on it, and then having to unwind it.
**What I am doing differently tomorrow.**
Running every East African query between six and nine in the morning, and again around five in the afternoon, never overnight.
Running the capture scan on every batch rather than every suspicious batch, because the suspicious ones are the ones I notice and the problem is the ones I do not.
Keeping the six surviving phrases and adding new ones only through the four desk tests.
Counting usable rows per batch in writing, because the count is what turned one vague problem into six specific ones, and I would not have found any of this by thinking harder about keywords.
And re reading my own detectors to check they do what I said they do. The one that mattered was two thirds written and described as finished, and it only got completed because I went back to make an earlier sentence true. That is not a process. That is luck, and I would rather it were a step.
If you are picking tools, vendors or sources on evidence you gathered this way, the evidence quality question comes before the comparison question. We built a short scorecard for exactly that decision, free and no email: https://t.co/30w9iH4meS
Ask me in a month what the seventh cause was.
You don't beat competitors by fixing what's broken.
You win by setting the bar on what great looks like,
then having the discipline to keep raising it.
Owning a business is one thing. Knowing how to sell what you offer is another.
You can have a great product, an excellent service, and a clear vision for your business, but without the right sales approach, it can be difficult to turn those into paying customers and ... (Thread)
earning trust, creating loyal customers, and developing networks that support long-term business growth.
Are you ready to build your business with the right mindset, skills, and strategy? Join the 2027 AVODA Cohort and take the next step in your entrepreneurial journey.