How to test what Claude Code (or Cursor, or Copilot) wrote before you merge it
Test an agent's change the way a customer would meet it: open the product, walk the journeys that matter, and check that each still ends where it should. Reading the diff tells you what the agent meant. Running the product tells you what it did. The check below takes about an hour to set up and then runs without you, on every push, before the merge.
Why reading the diff is not enough
An agent edits the files it was pointed at and whatever it decided to tidy on the way. The regression is usually in the second group: a helper renamed, a default changed, a redirect added. Unit tests on the edited function pass. The checkout page that used the helper does not, and nobody opened it.
Reviewing a model's diff line by line also does not scale. A team merging twenty agent changes a day cannot read twenty diffs with the care a regression needs. The review has to be a thing that runs.
The merge check, step by step
1. Decide what must keep working. Not everything: the five to fifteen journeys a customer would notice within an hour of a break. Sign in, the main create flow, search, checkout or its equivalent, the settings page people actually use. Write them as a list before you write any test.
2. Turn each journey into a test a browser can replay. Record it once, or let an exploring agent map the site and propose the journeys and review its plan. Each test should end in an assertion about an outcome, not a screenshot: the order number appears, the row is in the table, the URL is the dashboard.
3. Prove each test can fail. A test that passes on a broken page is worse than none, because it tells you the page works. Break the page, or let the tool do it for you, and confirm the test goes red. Only then does a green mean anything.
4. Run them on every change, from the pipeline. A deploy to a preview environment triggers the run with the environment's URL. The result comes back as a status the merge waits for, and a message to the channel when it fails, with the step that failed and what the page showed.
5. When a test fails on a page that changed on purpose, repair the test, not the product. Most failures after an agent's change are a moved button, a renamed label, a new step in a flow. A good tool proposes the one small change to the test, proves the repaired test on a fresh session, and waits for a person to approve it.
A worked example
A composed example: the team, the timings and the counts show the shape of the work and are not measurements from one customer.
A team runs a small SaaS with a billing page. They ask an agent to "simplify the plan picker". The agent merges a tidy change, and also removes a hidden input the invoice form read.
Their merge check has a test named "upgrade to Team": sign in, open Billing, pick Team, confirm, assert the plan badge reads Team and the invoice preview shows the new price. The run after the agent's push fails at the last step: the invoice preview shows the old price. The failure message names the step, shows the page, and links the run to the pull request. The agent is told, fixes the form, and the next run is green. A reviewer still reads the diff, and now reads it knowing which journey it broke and that the fix is proven. The check did not replace the review; it told the review where to look.
What to measure
Two numbers tell you whether the check is working. How often a run goes red on a change that would have reached customers, and how long a red takes to be resolved. If the first is zero for a month, your journeys are not the ones that break. If the second is more than a day, the failures are not readable enough.
FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.
See FlowQA