More and more teams work like this: a developer uses an AI agent to write the code, another AI tool reviews the pull request, and a human glances over it before merge. Three layers of review. It sounds safe.
And yet bugs still slip into production. Why? And what should testers do about it?
Why the bugs still get through
1. The reviewers share the same blind spots
Layers of review only help when their mistakes are independent. An AI that writes code and an AI that reviews it are often trained on similar data and share similar assumptions. If the writer misunderstands how an API behaves, the reviewer will often find the same wrong idea perfectly reasonable. Two reviewers who miss the same bug are no better than one.
2. Nobody sees the real intent
Most serious bugs are not typos. They are mismatches between what the code does and what the business actually needs. An AI reviewer sees the diff. It does not see the unwritten rules, the odd legacy behavior, or the decision made in a meeting last month.
3. Context is limited
Reviewers usually look at the changed lines, not the whole system. A change can be correct locally and still break something three modules away, or behave badly under real data, real load, or concurrency.
4. Clean-looking code lowers our guard
AI-generated code is tidy, well named, and confident. That makes it easy to skim. Add a green AI review on top, and humans trust it even more. This is automation bias, and it may be the biggest factor of all.
5. Everyone assumes someone else checked
The developer thinks the reviewers will catch it. The human reviewer thinks the AI already checked. The AI only comments on what it can see. Responsibility gets spread so thin that nobody holds it.
6. Big diffs get shallow reviews
AI agents can produce huge changes in minutes. Faced with 1,500 lines, a reviewer reads the summary, not every path. Subtle logic errors hide in exactly those places.
7. Some bugs only exist at runtime
Race conditions, performance cliffs, messy production data, config differences, and integration failures often cannot be seen by reading code at all, whether the reader is human or AI.
What testers should do
In this setup, the tester becomes more valuable, not less. Testers are the one layer that does not share the AI's assumptions. The job shifts from "does it run?" to "does it do the right thing?"
Test from requirements, not from code
Write test cases from the user story, business rules, and acceptance criteria before reading the implementation. Tests generated from the AI's own code tend to confirm what the code does, not what it should do.
Hunt where AI and reviewers miss
- Edge cases and boundaries: empty, null, very large, negative, special characters, unusual dates and time zones
- Error paths: timeouts, dropped connections, invalid input
- Permissions and security: can user A see user B's data?
- Concurrency: double clicks, two people editing the same record
- Integration points, since AI often breaks things outside the diff
- Older features the change might have touched
Take regression testing seriously
A change that is correct in isolation can still break something else. Keep a solid automated regression suite and run it on every change, not only on the new feature.
Use realistic data and environments
Many bugs only show up with production-like volume, messy real-world data, or different configuration. A staging environment that mirrors production is worth the effort.
Do exploratory testing
Scripted tests only check what you thought of in advance. Spend time using the feature like a confused, impatient, or malicious user. This is where the "nobody thought of that" bugs live.
Question the requirement itself
Ask what should happen when X occurs, and whether the behavior is really what the business wants. Testers often spot the gap between what was asked and what was needed, something no code review will catch.
Treat "AI reviewed it" as zero evidence
An AI approval says almost nothing about quality. Test as if no review happened. Push back on oversized changes and ask for smaller pieces that can be tested properly.
Be careful with AI-generated tests
AI is great for drafting test ideas and surfacing scenarios you forgot. But check that the tests assert meaningful outcomes rather than merely running without crashing, and that they do not copy the same flawed logic as the implementation.
Watch after release
Some bugs will only appear in production. Help set up logs, monitoring, and alerts, check them after each deployment, and make sure rollback is fast.
Report patterns, not just bugs
If the same kind of defect keeps appearing, such as missing validation or off-by-one errors, share that with the team. They can improve prompts, add lint rules, or add pipeline checks so the problem is caught earlier.
What the whole team can do
Testers should not carry this alone. Teams can help by:
- Keeping changes small and focused
- Using a different model or tool for review than for writing
- Having humans review for intent and risk instead of style
- Shipping with canary releases, feature flags, and quick rollback
The takeaway
AI has made writing code faster, but it has not made being wrong any less likely. It has only made wrong code look more convincing. Stacking AI reviews adds comfort more than certainty.
The safety net that works is independence: tests written from real requirements, humans who think about intent and risk, and a tester who refuses to assume that three green checkmarks mean the software is right.