TestFinch

Blog

Autonomous test generation: how it works, where it fails, what to review

2026-10-08 ยท An agent can map a site, propose journeys, write the tests and prove them. It still needs a person at three points. Here is what happens at each and what to look at.

Autonomous test generation works in four stages: an agent maps the site, proposes a plan of journeys, generates a test for each, and proves every test by breaking the page and watching the test fail. It fails in predictable places: pages it cannot reach, values it cannot know, and outcomes it cannot judge. The review it needs is small and specific, and it comes at three points: the plan, the questions, and the proof.

Stage one: map

The agent opens the site the way a person would, follows the navigation, and records what each page is for, what it contains, and where it leads. It stays inside the hosts you allow and stops at the depth you set. The output is a map: pages, their purposes, the forms and actions on them. A good tool keeps this map as notes the next run reads, so the second exploration starts where the first ended.

Where it fails. Pages behind a login it does not have, flows that need a real email or a payment, and anything that only appears after a background job. It should say so rather than guess.

Stage two: plan

From the map the agent proposes journeys: "sign up with email and reach the dashboard", "create a project and see it in the list". Each has a start page, the steps it expects, and the outcome that would prove it worked.

What to review. This is the first point where a person's time is well spent. Ten minutes on the plan beats an hour on the tests. Cross off journeys that do not matter, add the ones the agent missed because it could not see them, and correct outcomes that are wrong: the agent may think "a toast appears" proves a save when you know the toast appears even when the save fails.

Stage three: generate

For each journey the agent drives the browser, fills forms with values it invents or you provide, and writes a test from what it did, with assertions on the outcome.

Where it asks. It cannot know your test account's password, which of two similar buttons you meant, or whether it may follow a link to another host. A good tool asks, with a screenshot, and waits. The answers are kept, so it asks once. On a desktop it can also hand you the browser for the part it cannot do, and take the recording back when you are done.

Stage four: prove

A generated test that passes proves nothing yet; it might pass on a blank page. So the agent breaks the page, by removing the element the test depends on or changing the text it asserts, and runs the test again. If the test still passes, it is rewritten or thrown away. Only a test that fails on the broken page and passes on the real one is kept.

It also replays each test on a fresh session, because assertions written while logged in and warmed up often depend on state the first run will not have.

What to review. The proof results. A journey that could not be proven is the agent telling you it does not have a reliable assertion for that outcome, and that is where a person adds one.

A worked example

A composed example: the team, the timings and the counts show the shape of the work and are not measurements from one customer.

A team points the agent at their staging site with one test account. The map finds twenty-three pages. The plan proposes eleven journeys; the reviewer removes two (an admin-only report nobody uses) and corrects one outcome. During generation the agent asks three questions: a verification code from the inbox, which of two "New" buttons to use, and whether it may follow the link to the billing provider's hosted page (no). Nine tests are generated; eight are proven. The ninth asserted a date that changes daily, and the agent flagged it rather than keeping a test that would fail tomorrow. The reviewer fixes that assertion by hand. Total person time: about forty minutes, for nine tests that run on every deploy from then on.

The honest limits

The agent is as good as the site is navigable. A product that hides everything behind hover states and custom widgets with no labels will need more questions and more review. That is also a finding about the product, and often a cheaper one to fix than the tests.

FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.

See FlowQA