TestFinch

Blog

Regression testing for AI-generated code: a practical guide

2026-10-08 ยท When a model writes most of the diff, the review you used to rely on is gone. What takes its place, and how to set it up without a QA team.

Most teams that adopted coding agents noticed the same thing within a month: the volume of change went up and the number of people who had read each change went down. That is not a complaint about the models. A person who wrote a change by hand carried a mental model of what it touched; a person who accepted a change from an agent often did not. The review that caught regressions used to live in that mental model. It needs a new home.

What a regression is, in this world

A regression is a behaviour that worked yesterday and does not work today. It is rarely in the file the agent edited. It is in the checkout page that depended on a helper the agent "cleaned up", in the form whose validation moved, in the redirect that now goes one hop further. Unit tests on the changed function do not see any of this. The only test that sees it is the one that drives the product the way a person does.

The three properties a regression test needs

It has to run on a fresh session. A test recorded while you were signed in with your own account, with items already in a basket, passes on your machine and nowhere else. Every assertion has to hold on a browser that has never seen the site.

It has to be able to fail. A test that opens a page and checks the URL will pass while the page shows an error. Before you trust a test, break the page the way the test claims to protect it, and watch the test fail. If it does not, it protects nothing.

It has to run without a person. On deploy, on a schedule, on demand from a pipeline. A test that is only run when someone remembers is a test that is run on the day after the regression shipped.

Where the tests come from

Three sources, in the order most teams end up with them:

  1. Recording. Click through the flow once in a real browser and keep the steps. Fast, and good for the flows you know matter. The weakness is coverage: you only record what you thought of.
  2. Exploration. An agent maps the site, proposes journeys with a goal and a success condition each, and generates a test for every journey you approve. The weakness is judgement: it will propose journeys you do not care about, and it needs the two checks above to stop it generating tests that cannot fail.
  3. Bugs. Every bug that reached a person is a regression test waiting to be written: the steps that led there are the test, and the broken behaviour is the assertion. Teams that turn bugs into tests stop fixing the same thing twice.

A setup that works for a small team

None of this needs a QA team. It needs a place where the tests live, a way to run them without a person, and the discipline to reject tests that cannot fail.

FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.

See FlowQA