TestFinch

Blog

How to write a test plan a model can execute

2026-10-08 ยท A plan written for a person leaves out what the person already knows. A plan a model can run names the start, the steps, the data and the outcome, and nothing else.

A test plan a model can execute lists journeys, and each journey names four things: where it starts, what to do in order, what data to use, and what outcome proves it worked. Everything a human tester would infer from context has to be written down, and everything that is not one of those four things should be left out. The reward is a plan that turns into running tests in an afternoon.

What a person fills in that a model cannot

A human tester reading "check that checkout works" knows which product to pick, that the test card is 4242, that the order confirmation page is the outcome, and that the promo code field can be ignored. A model knows none of that unless the plan says so, and when the plan does not, it guesses. Guesses are where generated tests go wrong.

The four parts of a journey

Start. A path, and the state the browser is in. "Start at /projects, signed in as the test account." If the start needs a shared opening (sign in, pass a gate), name it.

Steps. Short, in order, each one an action a browser can take: open, click, fill, choose, press. Name elements by what a person sees: "the New project button", "the Name field". Do not name them by selector; that is the tool's job, and selectors in a plan rot.

Data. Every value the steps need: names, amounts, the test card, the email. Secrets by name only: "the password" means the project's secret called password. If a value must be unique each run, say so: "a project name with the run's timestamp".

Outcome. What on the page proves the journey worked. Be concrete: "the project appears in the list with its name", "the URL ends in /dashboard", "the total reads $19.00". "It works" is not an outcome. A toast that appears even on failure is not an outcome either.

What to leave out

Why the feature exists, who asked for it, the edge cases you are not testing, screenshots of the old design. A plan is not a specification. If a journey needs a paragraph of explanation, split it into two journeys.

Give the model the site's notes

A plan is executed against a site the model has to understand: what each page is for, which form does what, where the pitfalls are ("the Save button is disabled until the name is unique"). Keep those as notes on the project, written by the people who know the product and by the agent's own explorations, and corrected when they are wrong. A model reading the plan with the notes beside it asks fewer questions and guesses less.

A worked example

A composed example: the team, the timings and the counts show the shape of the work and are not measurements from one customer.

Before:

Test that users can invite teammates and that the invite works.

After:

Invite a teammate

  • Start: /settings/members, using "sign in".
  • Steps: click Invite people; fill Work emails with the invitee email; choose Member as the role; click Send invitations.
  • Data: invitee email is a unique address at the test inbox domain.
  • Outcome: the Pending invitations list shows the invitee email with the role Member.

Accept an invitation

  • Start: the invitation link from the test inbox, in a fresh session.
  • Steps: fill Name; click Join.
  • Data: a name; the invitee's password is the secret called password.
  • Outcome: the URL is /projects and the header shows the workspace name.

Two journeys, each with a start, steps, data and an outcome. A model can run both, and the tests it writes assert the outcomes the plan named rather than whatever it noticed on the page.

Reviewing the plan the model proposes

The same four parts are the review checklist when an agent proposes a plan from its own exploration. Does each journey start somewhere reachable? Are the steps things a browser can do? Is the data real or invented? Is the outcome something that would be absent if the feature broke? Fix the plan before generating, and the tests come out right the first time.

FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.

See FlowQA