Agile Coach & Scrum Master | Transforming teams with Agile, Scrum, Kanban, & SAFe | Passionate about driving efficiency & innovation in software development
The notion of an MVP has gone the way of Agile. It started out as a valuable, well-defined (and useful) technical term, but has been transmogrified into complete nonsense by people who have no idea what they're talking about, which doesn't stop them from talking about it 😄. (See "Dunning-Kruger Effect.")
The original term was coined by Frank Robinson c. 2000, but it was adopted (and changed) by Eric Ries in his book "Lean Startup." I use Reis's definition: the MVP is the smallest possible test of the viability of a product idea. It is not a product. It's a test. It tests the core of the core of the core of the core idea of a potential product. That's it. In this day and age of AI assistance and cloud services, if it takes longer than a couple of days to put together and deploy, it's way too big.
MVPs are often throwaways, but building them as real production code, complete with tests and a decent UX, is not a bad strategy. That will only add a day or so to the dev time, given the minuscule size we're talking about. That way, you can use it as the basis of a "walking skeleton."
If you release your MVP and people say "This is great! Where do I buy it?", you're done. Think of all the time and effort you just saved!
Usually, though, the reaction is "This is good, but…" That "but" tells you what to do next. Find out what you're missing, then modify your product to do the smallest, most minimal thing you can get away with to satisfy that "but." Then release it and repeat.
No bells and whistles. No big "features." No requirements gathering. No weeks of collecting something you call a user story, but probably isn't. (Actually, that "but…" is a user story, but that's a different post.) No big up-front spec that you feed to your AI. (Sure, use an AI, just not like that.) You need almost no process at all, in fact, so no Sprints and Sprint Reviews and Backlogs and all that other garbage. Build small, release often, get feedback, adjust.
What you now have is a small but viable product that's growing incrementally to meet actual expressed needs. You have something you can sell almost immediately.
And I don't want to hear, "In the real world of Niagara Megacorp, we can't do that." No. You can't. Your competitors can, however. (And yes, you can't build everything this way, but that's an uninteresting corner case.) Meta just went nuts because Zuch is terrified of exactly this problem and is reacting like the headless chicken he is. He should be scared.
If you are "writing a user story," you're already in trouble. A "user story" is exactly that, the user's story. They're the ones who "write" the story, not you.
A user story is not a code word for a waterfall up-front specification. A "story" describes your user's work, not yours. It does not describe a solution to a problem; it describes the problem you're solving.
When I say the above, I always get pushback from people who somehow imagine that I'm asking users to write essays. In other words, they're imagining a waterfall requirements document created by the user. Nobody writes anything of the sort. Not them. Not you. I put the word "write" in quotes for a reason.
"Stories" are conversations, not documents, certainly not tickets. Feel free to take notes, but a story is not a specification, and nobody literally writes it down. Have the conversation just before you implement, so it's fresh in your mind.
If the first time you get feedback is in your "Sprint Review" or "Demo," your Sprint has probably failed. Your goal is to deliver the most valuable thing. If the review feedback tells you that you need to make a change, then you have not delivered the most valuable thing.
You don't want to build the wrong thing, then waste time turning that wrong thing into the right thing. It's better to build the right thing to begin with. You needed to learn about that change _during_ development, when the change is easy, not after.
I should add that if the feedback you get in that "review" is from "stakeholders," not users or customers, then the failure window is multiple Sprints. You don't know if you're building the right thing until real users/customers see it. No customer surrogate will tell you that.
You do not build trust through estimation. You build trust by delivering. I've found that once I start delivering something small but useful every day or so, people stop asking for estimates. Managers no longer need estimates and milestones to monitor progress because they can see the progress with their own eyes. Every day.
Transparency trumps opaque process every time.
There is no predictive value to story points. A one-point story can take anywhere from a day to several weeks. A 13-point story can take anywhere from a day to several weeks.
For the most part, people who use points have never done the math. I suggest making a chart with points on one axis and actual measured development time (cycle time) on the other. You may be surprised by the results.
For 20 years software had a stable playbook. Need reliability? Here are five ways to write tests.
Now that writing code is nearly free, the engineers best at that playbook turn to it, and it's blank.
I've been reading posts saying AI puts an end to "Agile." That's nonsense.
AI agents are an irrelevance in an Agile context. Agents cannot tell you what to build, which is the key problem. It was never about coding, so much as writing the correct code.
The point of Agile is to build software that people actually want/need through continuous learning. We release small increments to users/customers, get their feedback, and adjust. It is about having the flexibility to improve both what we build and how we build it as we learn. Historically, it addressed the problem of big up-front plans, which were the pre-Agile norm, not working, at least if you define "working" as building stuff people actually want to use.
All that other garbage people cite (standups, sprints, &c.) has nothing whatever to do with Agile. The one exception is the retro, which is mentioned explicitly in the Agile Manifesto and is central to the learning process. You have read the Manifesto, haven't you? Doesn't seem like many of the "Agile" haters have. They seem to think that Agile is a process or framework. It isn't.
A backlog is a collection of work items that you do not have the capacity to build. (If you did have the capacity, you'd just build it.) It's a black hole where work goes to die. The only way to alleviate that logjam is to increase capacity, either by adding people (usually doesn't work) or by improving how you do the work (most companies won't). Consequently, the backlog always grows over time, filling with work that will never be done.
AI might fall under the "improve" category, but in practice, it doesn't seem to. My theory is that this is Parkinson's Law in action—work expands to fill all available capacity. If you add AI to the mix, the backlog grows faster than the amount of work removed. I've also seen people use AI for work not specified on the backlog at all—they just get carried away by the "productivity," adding unspecified feature after unspecified feature as they work. If you define "productivity" as removing backlog items, the AI gets you nowhere.
My solution is to get rid of the backlog altogether. Just throw it out. All of it. Now. If it's important, it will come back.
Then breathe a sigh of relief.
Then create, not a backlog, but a fixed-size input queue for Engineering. The size is just large enough so that, when you need something to do, there's something to do. A couple of weeks' work is plenty. Then chisel that size onto stone tablets and guard it zealously. Nothing goes on unless something comes off. Period. (This is called a work-in-progress, or WIP, limit.)
Nobody, not even the CEO, can violate the size limit. If some piece of do-it-last-week-or-the-sky-will-fall work comes along, you need to remove a queue item to make space. A discussion with whoever inserted that item will ensue. That's a good thing. The queue must hold the most valuable (to users/customers) work, and if everything is "most valuable," nothing is. The size limit guarantees high value.
#NoBacklog
"Monolith" is a deployment strategy, not an architecture. All it tells you is that you deploy the entire system in one go. It tells you nothing about the structure of that system.
"Microservice" is also a deployment strategy, not an architecture. Microservices allow you to deploy individual components (that's the architecture—autonomous components) without deploying the rest of the system at the same time.
You get what you measure:
Measuring commits encourages people to do unnecessary work. Commits don't improve quality or get the product out the door faster. When there are too many PRs, people stop reviewing them. It's easy to commit every couple of minutes, so I'll do that.
Measuring lines of code encourages people to create more code, even when that extra code adds complexity that will make everything harder moving forward. It leads to vast quantities of code nobody wants, code that's expensive to maintain and slows testing and development without increasing revenue. The best programs get smaller over time, not larger.
Measuring test coverage leads people to write tests that don't verify whether the code actually works or solves domain-level problems. As long as it's covered, who cares if it works?
Measuring hours spent typing (i.e., busyness) accelerates burnout, which, in turn, makes people work less effectively. Busyness also makes it harder to solve unexpected problems when they emerge (because you're too busy to get to it). It takes 30%-50% slack time to ensure a smooth, efficient development flow.
Measuring token use leads to undiscovered bugs and security vulnerabilities hiding in the output volume. It also leads to ineffective use of the AI and increased development costs.
Measurement is not harmless.
My first rule of testing is: don't change the tests.
To be more precise—write the tests in such a way that they don't have to change.
When you refactor, you test to make sure the code behaves as expected, change the code, then run the test again to make sure you haven't broken anything. If your tests change when the code changes, you cannot refactor safely. If the test doesn't exist, you have to write one before you start. In an AI world, every AI-induced change is a refactor, and often massive multi-unit refactoring is going on. No tests? No safety. Even with TDD, you add tests as the code evolves, but you don't change existing ones.
Among other things, this means you cannot safely test the minutiae of the implementation, because it changes constantly. Test behavior instead. "When I ask a component to do X, I expect it to do Y" ("Gherkin" format). The "unit" must conform to the Liskov Substitution Principle: clients don't care about even radical changes, provided that the interface (in the general sense, not the Java &c. sense) hasn't changed. In a microservice, that radical change could be a complete rewrite in a different language running on a different platform. The arguments to a method are its interface, for example, as are the APIs or messaging interface to a remote service or the set of public methods of a class, whether or not there's a formal "interface" involved.
Of course, that brings up the issue of what if the interface _does_ change?
In general, I find that constant changes to interfaces usually imply ad hoc programming, where nobody thinks things through before writing the code. People throw whatever pops into their heads out there and hope for the best, just like an AI does. The programmers are constantly surprised as they work, and are constantly rewriting to handle those surprises. That's coding without architectural thinking. Code like that is insanely difficult to test.
So, what's the alternative? Think about Domain-Driven Design. In DDD, every component, from single methods up to aggregate subsystems, models something in the actual domain. The name is a name you'll find in the domain. The messages and interfaces reflect a conversation that people would have if they were doing the work manually. The domain tends not to change, so interfaces that model the domain tend not to change. Note that this approach does not require a big-up-front design. The domain IS the architecture, so you can implement that architecture one piece at a time, and the domain itself guarantees architectural coherence.
More to the point, this is something you can test.
I recently came across this misguided poll 👇. The answer is none of the above. The biggest QA challenge is creating a culture of quality that pervades the organization. As Deming said, "Inspection does not improve the quality, nor guarantee quality. Inspection is too late. The quality, good or bad, is already in the product." If you don't find bugs until after development, your development process is broken.
No software development organization needs a standardized process.
40 years ago, Watts Humphrey, funded by the US Department of Defense, came up with something called the Capability Maturity Model (CMM)—one of many grasping-at-straw attempts to get DoD spending and quality under control. The model was forced onto all DoD contractors and eventually infected the corporate world. Everybody at the time was doing hard-core, huge-design-up-front, phase-gated (waterfall) software development, and CMM assumed that was just the way the world worked. CMM thinking mandated that the entire organization standardize on a single process under centralized control. It was the military, after all. Hierarchy and uniformity were a given.
Like SAFe nowadays, CMM was better than chaos, but didn't really achieve its objectives by most measures. Nonetheless, ideas from CMM are the corporate version of the zombie apocalypse. They refuse to die and define an of-course-things-have-to-work-that-way alternative reality.
For example, it's a CMM notion that the entire organization must march in lockstep to the Scrum (or SAFe, or (formal) Kanban Method, or lately, Spec-driven-AI) drum. We don't. The people best equipped to figure out how to do the work are the ones doing the work. It's a system based on trust, based on the idea that the people you hired because of their competence are actually competent, and you have to trust them to get the work done.
In the best organizations, the teams come up with their own processes. Sure, they can draw on the Scrums and Kanbans of the world for ideas, but ultimately, no standardized process works everywhere (anywhere?). Processes must be customized to the needs of the teams. This is not a new radical notion, by the way. Toyota has been doing it on its lines for 60 years.
(Sidebar: the people in the "don't understand Agile but nevertheless hate it " community claim that Agile advocated no process at all. That's nonsense. Process is good, but you need agility in your process thinking. Self-organizing teams figure out the best way for the team to work. They have a process, just not _your_ process.)
Of course, you can't have an every-team-for-itself chaos. The system is not adversarial, full of lazy people who hate one another, as many seem to think. The teams coordinate with each other constantly and adjust and adapt as needed. (Spotify calls this "alignment.") If they can't find alignment, management provides coordination and guidance (not directives and force) and always provides active support. Alignment is hard to do remotely, so get everybody together in one room every so often.
So, process is good. Trusting teams to figure out how best to do their own work is good. Whipping people into uniformity is not. Let's all start acting as if we're a community, not a gulag.
<rant> Story points are like cockroaches. Just when you think that you've stamped them out another army of them crawls out from under the sink.
Points have no value. None. Nada. They are an utter waste of time. Fibbonocci sequences just heap more garbage onto the pile. If you're working on the most valuable thing, I don't care how long it takes. You have no choice to not do it. Sure, work on narrowing stories to represent the smallest possible scope. Small batches are important. But "can we make this smaller" is a very different question than "how long will this take." Rather than wasting time on points, use that time to build.
And no, the discussion around points has no value, either. That discussion is focused on the wrong thing—implementation detail. The important discussion is on value to the customer, and that happens before, during, and after the work. It is not something you do only before you start working. That "can we make this smaller" discussion is about value—narrowing scope. Ask "do we really need this?" before writing every line of code. If you want to discuss something, discuss that.
</rant>
When programmers talk about how productive they are with AI, they tend to forget that they are working for businesses, where the only metric that matters is profit, and where the behavior of the overall system of work is the important thing, not the behavior of one small part of it.
Producing 3x the code using an AI does not benefit the business unless that extra volume increases revenue by more than the cost. If that extra code didn't make extra money, you've done nothing useful. If you use 1/3 the programmer time to create a unit of code but spend 3x the time down the line debugging, refactoring, and otherwise touching the code, you also have zero benefit.
It's easy for a programmer to imagine they're more productive if they're not factoring in the downstream costs of their work—work done by somebody else. Finally, in a complex system like this, no claim is valid unless you have actual full-lifecycle (not just development) metrics to back it up.
I am not saying there isn't a benefit or productivity increase associated with AI, but unless I see the numbers, I'm reluctant to believe those claims, particularly if they focus on one small corner of the system (e.g., a single programmer) instead of the system as a whole.
In my normal programming work, I use AI tools in a very focused way that I've been told is "wrong," even though I'm more productive, but there we are. I'm not spewing out 20x my previous output, so maybe that's it. I'm an engineer, so when I work with an LLM, I do engineering. Shocking, I know.
I work incrementally, in small batches, with a well-defined architecture guiding the process. (People talk nebulously about "guardrails." Having a coherent AI-friendly architecture is an important one.) I do not let the AI design the program for me. I build thin vertical slices with domain value, not random features. I focus on quality, not output volume.
I tell the LLM what the code should do, of course, but I also tell it how to structure the code—how it fits into the architecture, what the APIs or messaging looks like, etc. I don't simply describe the surface behavior of a "feature" and expect the AI to do all the work under the covers. Instead, I ask the AI to write small, manageable, focused, black-box testable components that fit into an overall system architecture that it does not control. The architecture comes first, and then I tell the LLM to implement it one component at a time. (I should add that some architectures work better than others in this context. DDD nested aggregates and entities have worked pretty well for me, as have message-based components with hardened perimeters.)
I'll admit that I often don't give the output more than a cursory glance, but I'm a fanatic about testing. (There's another "guardrail.") I don't let the AI write my tests. When I first started doing this, I found that the AI would modify the tests so that the incorrect code it generated would pass. That's not testing. Also, the AI-generated tests tested the wrong things—e.g., small implementation details, not domain-level behavior.
I'm a big TDD guy, so I write my tests first and then tell the AI that the code it creates must pass them. When I do use the LLM, I'll specify the test in Gherkin format, and the results are small enough to review manually. To me, manual review of the tests is essential. There's no room for AI ambiguity in a test.
I've been told that I'm not leveraging the full capabilities of the AI by working this way, that I should just describe features or modifications and have the AI do all the work. I am, nonetheless, more productive than when I don't use the LLM, and don't seem to have the problems (e.g., lurking bugs, fragility in the face of scaling, overwhelming complexity, etc.) that seem commonplace with other approaches.
My measure for productivity is time to complete a story, which has gone down. I couldn't care less about output volume. When working in the small like this, with the work constrained by component boundaries, the LLM cannot break existing code when it makes unrelated changes. I'm not overwhelmed by a pile of code so vast that I can't understand it.
So, maybe I am doing it "wrong," but I'm happy with what I'm doing.
My measure of productivity is the time it takes to get the solution to a user's problem (their story) into their hands. I couldn't care less about vanity metrics like output volume or token counts. What matters is delivering value sooner. I've found that AI helps with that, but if that's not the case with you, you might want to reconsider how you're using AI.
I talked about MVPs yesterday, but there are two related concepts, often confused with MVPs, that are also worth mentioning.
A prototype is a test that answers the question: “Does this design actually work the way we think it does?” You can prototype any aspect of a design, from the UX to the database. Prototypes come in two main flavors.
A low-fidelity prototype is something like a simple whiteboard sketch of a proposed UI. It goes together in seconds and is meant to help with quick decisions.
A high-fidelity prototype is a fully functional model that usually tests the behavior of some part of the program under load or in some real-world usage scenario. It can’t do that if it isn’t full-on production-quality code. Usually, high-fidelity prototypes are massaged over several iterations until they work as expected, and then are inserted into the main body of the code.
Neither of these prototype flavors is an MVP. A prototype tests a design. An MVP determines if a product is worth building. Those are entirely different things, even if they have a superficial similarity.
The other related idea is proof of concept (PoC). You use PoCs very early in the design process to prove that some aspect of the design will work. Typically, they’re used before you get to the prototyping stage and don't look anything at all like the finished product. For example, if you’re worried about database performance, you might kludge together a simple program that does nothing but dump data into the database and pull it back out to make sure the database can keep up with worst-case load. The concept is similar to the notion of a "spike" in Scrum. A PoC is neither an MVC nor a prototype. It’s just testing whether a design decision that’s still under consideration will work out.
So, we have three tools to work with:
An MVP tells us that a product is worth building.
Prototypes tell us that the design of the product is working.
PoCs tell us that core implementation assumptions are correct.
All three of these tools are essential, I think, in product development, but they are very different from one another.
All too often, the word "leader" is used to describe a mere manager. That's not a leader in any real sense. You cannot be anointed as a "leader" by upper management. Leadership cannot be imposed. A person in that position claiming they're a leader is puffery—a source of (often quiet because they're in a position of power) ridicule.
Leaders don't call themselves "leaders." They lead simply by being themselves.
Leadership is a personality trait. You cannot train a person to be a leader. The notion of a "leadership seminar" is a snake-oil rebranding of management training.
Teams confer leadership. It emerges when someone behaves in an inspiring way. A true leader does not have "followers" in the conventional sense. Mindlessly obeying or mimicking someone is cult behavior. Being forced to follow or obey is bullying, not leadership.
Leaders set examples, not impose ways of working. They make sacrifices, not require them. They inspire, not demand. They don't require respect; they're respected. Force and leadership cannot coexist. A power dynamic is a bully's tool, not a leader's.
So, I find that whole "leadership" framing to be disengenuous at best. I wish we'd drop it altogether.
One of the dysfunctions that AI amplifies is the software industry's focus on output. The move to AI is all about increasing output. Robot monkeys type faster and produce more output than people. The problem, of course, is that it's never been output that's the problem. Producing a garbage product faster benefits nobody; never has. It's good outcomes (e.g., happy customers, painless improvements) that we need, not more output.
That's not to say that, within limits, producing code faster is a bad thing. For one thing, getting working code into customers' hands faster gets us better feedback sooner. But. The best way to increase speed, whether or not AI is in the picture, is to not build things nobody wants. AI often does the opposite. I was reading this morning about the huge security holes in systems created by several of the vibe-coding platforms. Nobody wants security breaches. The customers certainly don't, and ultimately, given the legal liability and loss of customer goodwill, neither does the company. The same applies to even big platforms like Amazon, getting buggier and buggier by the minute. Nonetheless, the siren call of more output seems to push companies into doing things that are not in their best interests.
The other related assumption is that more code is somehow a good thing. That's also never been true. More code adds complexity. It hides bugs. It's harder to maintain (even with an AI doing the maintaining—loading 500K lines of code into a single context to fix a bug is not only expensive, but will probably introduce five more bugs for every one you fix). We always want the smallest, simplest thing that solves exactly the problem at hand and nothing else. Ego-driven development by people who pat themselves on the back and do a happy dance every time they create more more more, without bothering to look at what, exactly, that "more" entails, gets us nowhere good.
I have to add—to head of the inevitable wild-eyed cultists—that I am far from opposed to AI assistance. I use it myself, and any software shop that doesn't avail itself of effective tools has got worse problems than an output focus. However, we have to use AI in the context of producing value; value to both the customer and to the engineers. Value is not proportional to quantity.