Regression testing for AI-generated code: a practical guide
Most teams that adopted coding agents noticed the same thing within a month: the volume of change went up and the number of people who had read each change went down. That is not a complaint about the models. A person who wrote a change by hand carried a mental model of what it touched; a person who accepted a change from an agent often did not. The review that caught regressions used to live in that mental model. It needs a new home.
What a regression is, in this world
A regression is a behaviour that worked yesterday and does not work today. It is rarely in the file the agent edited. It is in the checkout page that depended on a helper the agent "cleaned up", in the form whose validation moved, in the redirect that now goes one hop further. Unit tests on the changed function do not see any of this. The only test that sees it is the one that drives the product the way a person does.
The three properties a regression test needs
It has to run on a fresh session. A test recorded while you were signed in with your own account, with items already in a basket, passes on your machine and nowhere else. Every assertion has to hold on a browser that has never seen the site.
It has to be able to fail. A test that opens a page and checks the URL will pass while the page shows an error. Before you trust a test, break the page the way the test claims to protect it, and watch the test fail. If it does not, it protects nothing.
It has to run without a person. On deploy, on a schedule, on demand from a pipeline. A test that is only run when someone remembers is a test that is run on the day after the regression shipped.
Where the tests come from
Three sources, in the order most teams end up with them:
- Recording. Click through the flow once in a real browser and keep the steps. Fast, and good for the flows you know matter. The weakness is coverage: you only record what you thought of.
- Exploration. An agent maps the site, proposes journeys with a goal and a success condition each, and generates a test for every journey you approve. The weakness is judgement: it will propose journeys you do not care about, and it needs the two checks above to stop it generating tests that cannot fail.
- Bugs. Every bug that reached a person is a regression test waiting to be written: the steps that led there are the test, and the broken behaviour is the assertion. Teams that turn bugs into tests stop fixing the same thing twice.
A setup that works for a small team
- Point an exploring agent at staging. Review the plan it proposes: drop the journeys that would place real orders or send real mail, keep the ones that cover money, sign-in and the main path.
- Accept only tests that passed twice on fresh sessions and failed on a broken page. Expect a third of what the agent proposes to be rejected; that is the check working.
- Run the suite from your deploy pipeline with a token, and nightly on a schedule. Alert the channel where the people who can fix things already are.
- When a test fails on a changed page, decide whether the page changed on purpose. If it did, let the agent propose the repair and approve it; if it did not, you have found the regression before a customer did.
- File what slipped through as a bug with the console and the network attached, and turn the bug into a test.
None of this needs a QA team. It needs a place where the tests live, a way to run them without a person, and the discipline to reject tests that cannot fail.
FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.
See FlowQA