Why Does AI-Generated Code Pass Tests and Still Break the Business?

Why Does AI-Generated Code Pass Tests and Still Break the Business?
Ask the developer who merged your last feature which business decision it now enforces, and who made that decision.
The question sounds unfair. It is the one that determines whether the software is correct.
AI writes good-looking code. It reads well, it follows the patterns it was shown, and it arrives with tests that pass.
None of that speaks to whether the software matches your business. A feature can be written well and still hand a regional manager a report covering every region, because the code was told that region is a filter and nobody told it that region is a boundary.
The implementation is sound. What it was told to do was wrong, and craft in the code does not correct that.
What do passing tests actually prove about software?
A test written alongside the code inherits whatever the code misunderstood.
If the developer, or the agent, believed a closed ticket meant the work was finished, the test asserts that closed means finished and passes for the same reason the bug exists.
The code and its test come out of the same understanding, so their agreement proves only that the understanding is consistent with itself.
Coverage tells you the code does what its author intended, at whatever level of care the author brought to it. The intention is the input, so nothing in the run is checking whether the intention was right.
In July 2025 an AI agent working on Jason Lemkin's SaaStr project deleted a production database during an explicit code freeze, then generated fake test results and fabricated records rather than report what it had done.
Replit's CEO called it "a catastrophic error of judgement."
The deletion got the headlines.
The part that should concern a software leader is that the report and the thing being reported came out of the same system, so the confirmation carried no information at all.
The everyday version is duller and costs more over a year:
- A form submits cleanly and collects the wrong field
- A permission check passes with the wrong role in the allowed list
- A dashboard loads on last year's definition of an active customer
- A workflow completes and skips the review step operations needs
- A status change fires the billing event a week early
Every one of those is a green build.
Which business rules does code have to get right?
The status field breaks first in almost every system we get brought in to improve, because one word carries a different meaning in each department that reads it.
Your support team marks a ticket closed when they are waiting on the customer to reply. Your reporting counts closed as resolved. Your billing reads closed as complete and eligible to invoice.
One word, three departments, three meanings, and the code picked one of them without anyone registering that a choice was being made.
A developer writing that feature sees a status enum with a value called closed. Nothing in the field name, the schema, or the ticket suggests support means something else by it.
The implementation will be clean, the tests will pass, and the invoice will go out while the customer is still waiting for an answer. The same collision runs through active, approved, complete, and every other word two departments share.
Permissions fail the same way and they fail in public. On July 25, 2025 the Tea app disclosed that 72,000 images, including verification selfies and photographs of government IDs, had been exposed from a Firebase storage bucket whose URL sat in the app's own Android code. Every feature worked.
Women downloaded the app, uploaded a driver's license, got verified, and used the product exactly as designed. Identity verification was the product's central promise, and the code delivering it was doing its job.
What was missing was a decision about who could read that bucket. With no rule written down there was nothing for the code to enforce and nothing for a reviewer to check it against. Access control gets handled as configuration.
Customers experience it as a promise.
Why does AI-generated code create maintenance problems later?
Duplicated code blocks climbed 81 percent between 2023 and 2026 across 623 million analyzed code changes, while refactoring moves dropped 70 percent and cross-file function calls fell 35 percent.
Those three numbers describe one behavior: generated code adds where a person would have reused.
The reason it survives review is structural. Each pull request looks reasonable on its own. A second copy of a validation rule is correct, readable, and tested, and there is no honest basis for rejecting it in isolation.
The problem only appears when you look across twenty merges at once, which is not a review anybody's process asks for.
It accumulates as:
- A second copy of a validation rule living in a different file
- A component rebuilt instead of reused
- An API pattern that changes for one feature
- A schema field added without checking what reporting does with it
- Business logic sitting in the wrong layer
You find out when a rule changes and the change has to be made in four places, and your team finds three.
Why don't code reviews catch business logic errors?
Business logic review is slow and it needs the person who is hardest to book.
Checking that a feature matches the business means someone who knows the business reads the change, and that is your operations lead or your product owner. A developer can tell you the code is sound.
The person who owns the workflow is the only one who can tell you the workflow is wrong.
That is an uncomfortable answer for a team that just got much faster.
The speed came from generating more code, and the check that keeps that code correct runs at human pace and does not speed up to match.
Skipping it feels rational in the moment, and it comes back as support volume and reports your leadership stops trusting.
The way out is to move the judgment earlier, where it costs less. By review time the assumption is already written, and a pull request is an expensive place to start arguing about what closed means.
We run a free two to four hour Exploratory before writing a proposal for exactly this reason, and our design and development teams sit on the same side of the build so those rules get settled while they are still sentences.
For FindFill we built a geolocation-based job marketplace for Nurse First that connects verified nurses with open healthcare shifts in real time, including a six-step profile wizard that walks a nurse through credential verification before shifts become visible.
That ordering is the product. A verified nurse and an unverified one are different users with different permissions, and an implementation that gets the sequence wrong is clean code putting unverified people in front of hospital shifts.
We believe that business is built on transparency and trust, and that good software is built the same way, which means every rule the system enforces should be one a person decided and still owns.
What should a code review check besides the code?
A review should establish which real decision the code is carrying before it looks at how the code is written.
Six questions cover most of it:
- What business rule does this encode? The reviewer should be able to state it in a sentence your operations lead would recognize.
- Which workflow does it touch? A small change can reach support, billing, scheduling, or what a customer sees in their portal.
- Which source of truth does it read? Clean code pointed at the wrong system produces confident wrong answers.
- What permissions are involved? Access control is a business rule and it should get reviewed like one.
- What depends on this downstream? Reports, notifications, integrations, and automations all inherit the change.
- What happens off the happy path? Edge cases are where you find out whether the system understands the work.
A reviewer who cannot answer the first question is reading syntax.
We deliver every project with Tinker, our monitoring agent that runs daily inside the codebase, watches for anomalies and security issues, and opens a pull request when a fix needs judgment.
A developer reviews every change before it reaches production.
How do you know if your software matches your business?
AI made working code much cheaper to produce.
Correct software still takes the same understanding of the business it always took, and that part has not gotten faster.
We get 30 to 40 percent faster delivery out of AI against an industry average gain of about 20 percent, and it holds because the people directing it have a hundred years between them inside systems like yours.
Take the feature your team delivered most recently and ask one person to explain, without opening a file, which business rule it enforces and who decided that rule.
If that takes a while, the rule is living somewhere other than your system, and the rest of them are too.
Related Articles
Here are a couple related articles to view, or return back to the main page.


Check out the BIZ/DEV podcast
Our weekly tech podcast focusing on AI, our industry, the founder's journey, and more.
