TestFinch

Blog

What a failing step looks like when a coding agent wrote the test

2026-10-10 ยท The recording on our homepage, step by step: three checkout tests written through the TestFinch connector, one outdated button name, the screenshot that showed the right one, and the rerun that passed.

The video on our homepage is a recording of the real app, the real workbench and a real FlowQA worker, run by one script on one machine against a small demo shop. The coding agent in it is scripted, so the story plays the same way every time; the run, the failure and the screenshot are real. This post walks through what happened, with the numbers the script printed.

A scripted coding-agent session: three checkout tests are written, one fails on a wrong button name, its screenshot shows the real button, and a fixed revision passes on the rerun. Sample data; the first run plays 5 times faster and the rest in real time.

The setup

The shop is Acme Supply: a catalogue of four products, a product page, a cart with quantities and a promo code, a checkout form and a confirmation page. Every button and field has an accessible name, so a test can find each one by its role and its label. The person is Mia Chen, and her project, Demo shop, points its tests at the shop. The script sends the same requests the TestFinch connector sends for a coding agent, announcing it as Claude Code, so every saved test says who asked for it and what wrote it, as it would in a real session.

The three tests

Three tests were saved in one call, which took 44 ms:

The card number is test data held in the test's own data, bound as {{card_number}}, never typed into a step. The save came back with no findings. In the app the tests appeared under Tests with the line "Written by Mia Chen through Claude Code".

The first run

One run started with all three tests on a cloud browser. It took 13.4 s from start to finish. The notebook test passed in 0.9 s and the promo code test in 1.0 s. The order test failed at step 9 after 10.9 s, almost all of it the ten seconds the runner gives a missing element to appear. The error was plain: no element for button "Pay now".

What the screenshot showed

The test named the checkout button "Pay now", an outdated name the script keeps on purpose, the kind an agent working from an old design would write. The shop calls it "Place order". Nothing in the product was broken; the test was wrong.

That is exactly the question a failing step's screenshot answers in one look. FlowQA keeps the screenshot of the step that failed, and opening it from the test's run history showed the checkout page as the browser saw it: the three fields filled in, and under them a button reading "Place order". A person reads that in a second, and an agent can read the same image through the connector.

The fix and the rerun

The test was read back, its one step changed to tap "Place order", and saved as revision 2 with the revision it was read at as its base, in 36 ms. A save with a stale base is refused, so a fix can never quietly overwrite someone else's edit. The header now read "Revision 2 by Mia Chen through Claude Code".

The rerun took 2.1 s, and the test passed in 1.1 s.

About the recording

The script sends the same requests the TestFinch connector sends for a coding agent, including a deliberately outdated button name, so the story plays the same way every time. The app, the worker, the run, the failure and the screenshot are real. The published cut is 41.5 s. The first run plays 5 times faster; everything else is real time. The script, pnpm --filter @snag/suite-site record-demo, seeds the data, plays the story in a headless browser, and fails if any step does not happen the way the story says, so the video can be recorded again whenever the product changes.

To try it with your own agent, connect it as the coding agents guide describes; the authoring guide is what it reads before it writes a test.

FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.

See FlowQA