VP for SQA²
QA leader with 10+ years' experience in automation, data-driven QA, and leading teams. Expertise in e-commerce, healthcare, media, and blockchain.
Took two of our team members, Brian Phan and Cody Lim, out to dinner last night at my favorite shabu-shabu spot.
We talked about their current projects, how it's been going, and how difficult it is to start something new. Starting something new is rarely easy. It's always challenging.
I shared my own experiences of taking on new challenges and growing through them. And I told them something I believe strongly. If you feel uncomfortable taking on a new challenge, that's usually a good sign. It means you're doing something you haven't done before. And that's where growth happens.
Growth never stops, no matter what level you're at.
The easy path is staying with what you already know. Nobody grows there.
Mentoring isn't about having all the answers. It's about sharing your experience so others can learn from it. Sometimes it's just dinner and an honest conversation about where someone wants to go and what's holding them back.
This is a common pattern I've seen across teams.
On paper, the process works. An issue gets caught before release. It gets documented. It gets passed to the right people.
It still makes it to production.
Not because testing missed something. Because the process depended on a handoff, and handoffs are where issues go to die. Everyone assumes someone else has it. The information exists. The ownership doesn't.
That's not a people problem. If a process can fail because one handoff didn't land, the process was designed with a single point of failure. People get busy. Things slip. A resilient process expects that.
This is the part of quality nobody wants to talk about. You can have the right gates, the right tools, the right documentation. If ownership isn't explicit at every step, the process is just paperwork.
And this matters more now than ever. AI is getting really good at the mechanical side of QA. Catching failures, sorting them, flagging what needs attention. That part is getting automated fast.
But AI can't own an outcome. It can't decide what's an acceptable risk. It can't make the call to hold a release. Those are human responsibilities, and they always will be.
So as more of QA gets automated, the human job doesn't shrink. It concentrates. Every person left in the process carries more weight, not less.
Build processes where ownership is explicit. Where handoffs are confirmed, not assumed. Where an issue can't die in someone's queue.
The tools will keep getting smarter. Accountability is still on us.
Had dinner last week with Luis Diaz and an old colleague and friend, Siamak Sadralodabai.
Luis and I go way back and still work together today. We met Sam in 2016 on the client side of a project all three of us were on. We stopped working together in 2020.
The relationship we built at a client evolved into a friendship. We've stayed in touch over the years and been there for each other outside of work too. We always have some friendly banter but the main thing is the friendship.
None of this happens by accident. Somebody has to send the text. Somebody has to set up the dinner. It takes being intentional about it. And when you put in that effort over the years, some of those relationships turn into something worthwhile.
Grateful to be able to keep in touch with old colleagues.
I was speaking with an engineering leader last week and he brought up a problem I see all the time.
If the people testing the product are the same people responsible for hitting the ship date, quality is competing with the deadline. Even on good teams.
But the deadline pressure isn't even the biggest issue. The bigger issue is blind spots.
Developers should test their own code. That's necessary. It's just not sufficient. When you test something you built, you test it the way you built it. You check the paths you thought about. The bugs live in the paths you didn't.
Some teams will say they have devs testing and a strong process around it, and that can work. Culture can patch the incentive problem. But structure beats willpower. An incentive you don't have to fight is better than one you fight well.
And to be clear, separation done wrong is worse than nothing. If QA is a wall at the end of the process that code gets thrown over, you've just added a bottleneck. The goal is an independent quality function that shifts left, embedded early in the process. Reviewing requirements before code is written, catching gaps in design, building automation into the pipeline. Preventing bugs, not just finding them.
Developers test their work. Someone whose only job is quality tests the product. You need both.
How does your team handle this?
I sat down with Kyle B., who's a senior product manager at Disney. We talked about how quality looks from his perspective on the product side.
A few things stood out to me.
He doesn't think about quality as bugs. He thinks about it as user expectations. If a feature works but doesn't do what the user expects, that's still a quality problem.
He told me about a QA analyst on his team who asked if a new feature would also work on Apple Vision Pro and Meta Quest. He hadn't even thought about those devices. One question from QA changed the scope before any code was written. This is what shift-left actually looks like. Quality doesn't start at testing. It starts at the requirements.
And that's where most problems come from. Requirements that are unclear, incomplete, or not testable. If QA only shows up after development, all they can do is find problems. If QA is involved when requirements are being written, they prevent them.
He also said something that stuck with me. Testing one feature is the easy part. The problems happen when multiple features ship at the same time and the teams building them aren't talking to each other. Quality breaks down between teams, not inside them. We see this with clients all the time.
Engineering owns the testing. Product owns the expectation. The best QA teams connect the two.
Automation tools have gotten really good. But I think it's worth being clear about what they're actually good at.
What tools do well:
- Execute tests at scale
- Ee-run on every commit
- Catch regressions in known scenarios
- Report results fast
What tools can't do:
- Identify what's missing
- Read intent from requirements
- Score whether coverage matches risk
- Decide which gaps matter most
Look at the two lists. The first is execution. The second is judgment.
Every item on the second list requires someone to understand the business, the users, and what failure actually costs. No tool knows that a missing authorization check matters more than a cosmetic bug. Someone has to.
This is why more automation hasn't closed the coverage gap for most teams. They bought execution and expected it to do judgment's job.
We call this Intelligent Quality. AI accelerates execution, not judgment. The teams getting real coverage are the ones who kept judgment in charge of what gets tested, and let the tools do what tools do well.
Had a great conversation the other week with Jason Lee, a senior product designer at Discord, about product quality.
We came at it from different ends. He's in product design, I'm in QA. That's what made it a good conversation.
For him, quality is the thing you don't notice. When it's done well, it feels seamless. You only notice it when something feels off. Does it look right? Does it work the way you expect?
I look at it from further upstream. For me, quality gets decided long before a user opens the app. It comes down to whether everyone actually agreed on what they were building. Most of what goes wrong doesn't start as bad code. It starts as a vague requirement, or two teams reading the same line two different ways.
What got interesting is we landed in the same place. Jason sees it too. When a requirement is unclear, it doesn't just slow things down, it shows up in the final product. The work is off, the direction is off, and the user ends up feeling it.
What he's designing for sits on top of the thing my side worries about. If the team never agreed on what they were building, it's never going to feel seamless to the user.
Someone asking, "does this feel right?", and someone asking, "did we build what we meant to?" You need both. Miss one and the user feels it.
I'm curious how others think about this one. Where does quality actually get decided for you?
Here's a pattern I see in almost every test suite.
The positive conditions are covered well. Can a user log in? Can they complete a purchase? Can they update their information? Can they run a report? Most testing lives here, because these scenarios come straight from the requirements.
The negative conditions are a different story. Can a user access another account's data? Can they complete an action without proper authorization? Can they bypass a required approval? Can the system quietly operate outside a mandated limit without alerting anyone?
These are the scenarios that end up in incident reports. And they're frequently missing, untested, and unautomated. Not because teams don't care. Because nobody wrote them down, and automation only runs what's written down.
Where testing concentrates is rarely where risk concentrates.
Next time you review your suite, don't count the tests. Ask how many of them describe what users should NOT be able to do. That ratio tells you more about your real coverage than the total ever will.
Caught up with one of our team members, Cody Lim, this week about one of our clients he's supporting.
This client is all-in on AI. It's built into how their whole team works, and they move fast because of it. That changes what supporting them looks like.
Our job isn't to slow that down. It's to make sure quality keeps up with the speed. So we've put specific workflows and guardrails around how AI gets used in the quality work. AI helps us move faster on execution, but nothing skips human review, and the judgment calls stay with our engineers.
That's what supporting a client actually means to me. You meet them where they are. We're not there to be a blocker. We call out risks when we see them, put guardrails in place, and improve the process where it needs it. If they're moving fast, our job is to make sure they can do it safely.
When I talk to engineering leaders, one of the pain points that keeps coming up is coverage. More tests than ever, automation running on every commit, and the same defects still showing up in production.
Here's why that happens.
Most automation re-runs what you've already identified. Every test in your suite exists because someone thought of it first. So your automation investment keeps making the known scenarios faster to check.
But the defects reaching production aren't coming from the known scenarios. They're coming from the ones nobody wrote a test for because nobody thought of them. Negative cases. Business rule violations. Abuse paths. Edge conditions.
Automation can't catch what was never identified. So the gap survives years of automation investment, because the tools were never pointed at it.
That's why we treat discovery as its own step. Before anything gets automated, we have a process for systematically identifying what should be tested, including the scenarios nobody thought to write down. Then we score the coverage and automate what matters.
The teams with the best coverage aren't the ones with the most tests. They're the ones who found the gaps before production did.
Caught up with Steven H. last week, a former colleague and client.
We met in 2014 when he was my dev manager. He's one of the best leaders I've worked with. He knew how to bring a team together and get things done. Back then we saw each other every day. Standups every morning, giving him my updates, walking through blockers. And lunches where the conversations went beyond work.
When he moved on to a new company in 2019, he brought my team in to support him there.
We've kept in touch ever since. Texts here and there, mostly about basketball, and we meet up when we can. We even caught a game together. Last week we talked about how our careers have evolved and how our roles have changed over the years. We talked about how AI has transformed our QA process, working with clients, personal growth, and how we're both trying to eat healthier these days.
One thing that stood out during our conversation was getting out of your comfort zone and taking on new challenges. Even someone as accomplished as Steven is still willing to do that. That's a lesson I'm taking with me.
We don't work together anymore, but I'm still learning from him. He still gives me advice I actually use. And we still talk basketball.
Everyone talks about networking. I just try to keep in touch with good people.
Thanks, Steven!
One of our clients build healthcare AI products that integrate with Epic. Our team tests them.
Everyone focuses on the AI. But when these products fail, the failure is usually in the integration, not the AI.
Most healthcare AI follows the same pattern. The AI generates something fast. A note, a summary, a recommendation. Then that output has to land in Epic, because Epic is the system of record.
The handoff into Epic is where most of our testing happens.
We check that the feature behaves inside the clinician's real workflow in Epic.
We check that the output makes it across the integration.
And we check that it lands in the record complete and correct.
So we never stop at the screen where the output was created. We verify it again inside Epic itself. If the output looks right in the app but lands wrong in the chart, the clinician is the one who deals with it.
AI speeds up the generation. It does not confirm the output reached the record.
Someone still has to decide what gets verified, where, and against what. That is judgment, and no model does it for you.
When Robert, one of our senior QA engineers, first talked to me about building an AI agent for one of our clients, we agreed on one thing before any code was written: no matter how much it automates, human judgment stays in the loop.
Last week I sat down with him to see where it ended up.
The scope is deliberate. Small, routine mobile bug fixes. Not their dev process. The problem was these fixes were eating their engineers' time. Senior developers working on bug fixes instead of new development. And every hand-off was improvised, so a fix could sit unverified until someone produced a build for weekly regression.
Here's how it works now.
An engineer tags the agent on a bug ticket and picks how much to hand off. Within seconds it acknowledges, posts a plan, and starts working in a sandbox.
At its lightest, it comes back with a root-cause analysis before anyone commits time. Go further and it implements the fix, pushes a branch, triggers a real device build, and hands off an install link with a step-by-step QA test plan. Its newest mode writes a regression test for that specific bug and runs it in CI with the full smoke suite. A passing run doesn't mean it's done. It means the evidence is ready for a human to review.
So every fix now arrives the same way. Structured hand-off, branch, build, test plan, proof. Nothing sits unverified anymore.
Their team is using it now and has seen major efficiency gains. It's also smart about failures. It can tell whether a test failed because of its fix or because of something else, like a flaky pipeline or a test that was already broken. And when it hits its retry caps, it escalates to a person instead of guessing.
But here's what we held back, on purpose.
It doesn't merge code. It doesn't open pull requests. It can prove a test passed. It can't tell you it fixed the right thing. That call stays with human judgment.
Because AI fails in ways guardrails won't catch. It states made-up behavior like fact. It writes a fix that looks clean but misses the real problem. Guardrails don't catch those. Judgment does. And judgment is the part we didn't automate.
That's what we call Intelligent Quality. AI accelerates execution, not judgment. It's what we build for the teams we work with.
Nice work, Robert.
AI has made execution cheap in QA. Test cases in seconds. Automation written faster than we can review it. The execution problem is basically solved.
But execution was never the whole job.
The other half is judgment. Knowing why a test exists. Knowing who stands behind a decision. Knowing what was actually verified. AI generates the artifacts. It can't answer for them.
So as automation scales, there are five abilities good quality engineering teams preserve. Each one is a question only judgment can answer:
1. Traceability: Can we connect risk, requirements, tests, automation, and results?
2. Accountability: Who reviewed and approved the decision?
3. Maintainability: Can future teams support and extend what we built?
4. Explainability: Can we explain why a test exists?
5. Auditability: Can we prove what was tested and when?
Every one of these has a person in it. Someone who connects, reviews, explains, proves. You can't generate your way to any of them.
If your team can produce ten thousand tests but can't tell you why any of them exist, you don't have coverage. You have noise.
We call this Intelligent Quality. AI accelerates execution, not judgment. These five abilities are how judgment keeps up.
What's your team preserving?
We had three new team members join our team recently. Whenever we have new team members join the team, I usually take them out to lunch.
It gives me a chance to get to know them outside of work and start building some camaraderie early. This time around I found out all three of them are gamers. I'm not much of a gamer myself, but a lot of our team is, so they already have something in common with the rest of the org before they've even settled in.
These lunches are also where I get to pass down some of the experience and knowledge I've built up over the years and do a little mentoring along the way. I've had mentors who did a lot for me in my career, and I try to do the same for the people coming up now.
I've always believed leadership should be available and personable with their team, not off in a corner somewhere. And teamwork is one of our core values, which to me starts with actually knowing the people you work with. A lunch is a simple way to do that.
Glad to have them on the team!
Every Monday morning I sit down with some of the senior QA engineers we've embedded inside our clients' teams. The agenda is always the same: our clients, and how we're showing up for them this week.
Not status updates. The real questions. Where can we take more ownership instead of waiting to be asked? Which quality engineering work should we be driving, not just running? What do the metrics say, and where are they hiding a risk we haven't named yet?
Because these team members live inside the work every day, they see the real picture, not a status report passed up the chain. And we never treat a good week as the finish line. Every release teaches us something about where to sharpen next. We plan, we deliver, we check the results honestly, and we adjust. Then we do it again, and our clients come out a little stronger every cycle.
In fintech, rework doesn't just cost hours. It costs trust.
When a defect reaches production in a payment flow or compliance-sensitive feature, the downstream effect is immediate: incident response, regression testing, possible regulatory review, and engineers pulled from roadmap work mid-sprint. That's not a QA problem. That's a sequencing problem.
The shift I push for with every fintech team is moving QA earlier, not just faster. When QA is embedded in the sprint and reviewing requirements before a line of code is written, defects get caught at the source. The rework becomes a 30-minute fix, not a 3-day hotfix cycle.
A pattern I see often: teams that invest in shift-left QA consistently reduce the volume of late-cycle defects, which directly compresses rework hours. One fintech team we worked with had a recurring issue where edge cases in their transaction reconciliation logic weren't caught until UAT. After embedding QA earlier in the development cycle, those same defect types were surfacing in code review, not in a production incident.
The fix was smaller. The blast radius was zero.
Developer hours are finite. In fintech especially, the question isn't whether QA costs time. It's whether you pay for it early or late.
Paying late is always more expensive.
AI is writing code faster than most QA processes can keep up with. This is a real problem.
When developers ship significantly higher volumes of code using AI assistants, test suites built for a slower pace of development start showing cracks. Edge cases get missed. Negative test paths go untested. Coverage looks fine on paper until something breaks in production.
The pattern is consistent: teams adopt AI coding tools, velocity jumps, and then QA becomes the bottleneck, or worse, the silent failure point.
The way I approach it is through an internal tool we use to systematically surface gaps, edge cases, and negative test scenarios, with AI helping identify what is not being tested rather than just what is. It shifts the focus from "are we running tests?" to "are we testing the right things?"
The cost difference between catching a bug in QA versus production is substantial. As code volume increases, that gap compounds fast.
How are you handling test coverage as AI accelerates your output?
Ask most engineering leaders how much time their team spends fixing production bugs, and you'll get a pause. Then a number that surprises them.
For many teams, it's a significant chunk of every sprint. Not building. Not shipping. Fixing.
That's not a staffing problem. That's a process problem.
When QA is treated as a final gate rather than an embedded practice, bugs accumulate quietly. They surface in production at the worst possible moment, and your best engineers get pulled off roadmap work to do triage.
The shift that actually changes this is moving quality left. When a QA engineer joins sprint planning and reviews acceptance criteria before a single line of code is written, defects are caught at the cheapest possible point in the cycle. That one step alone compresses the feedback loop significantly.
A common pattern I see: teams that embed QA from sprint kickoff, rather than bolting it on at the end, reduce the volume of production escapes over time. Not because they hired more developers, but because they stopped treating testing as an afterthought.
Engineering time is your most expensive resource. Spending a meaningful portion of every sprint on rework is not inevitable. It's a cost you can measure and reduce.
The question worth asking isn't "can we afford QA?" It's "what is production debugging actually costing us right now?"
Enterprise procurement teams are adding it to RFPs now. Same-timezone QA support, listed right alongside SOC 2 compliance and SLA requirements.
That shift is worth paying attention to.
When a critical issue surfaces in staging at 3pm on a release day, the response window is measured in minutes, not hours. I've watched teams lose entire release cycles because their QA partner was offline during the exact hours the engineering team needed answers. The async back-and-forth adds up fast, and the cost shows up in delayed releases and frustrated engineers.
The approach I've built around is simple: your QA team should operate on your clock. Same standup hours. Same Slack channels. Same incident response timeline. When something breaks, the conversation happens in real time, not the next morning.
What I'm seeing now is that enterprise buyers are formalizing this expectation. Timezone coverage is being written into vendor contracts, not just discussed in discovery calls. Procurement teams have started treating it as a due diligence item, the same way they'd evaluate backup and recovery procedures.
If your QA partner can't join a 2pm war room when production is down, that's a gap, and enterprise procurement teams are starting to price that risk accordingly.