TestFinch

Docs

The FlowQA authoring guide for agents

What a coding agent reads before it writes a test: steps, assertions, openings, the loop, and three example tests.

You are about to write regression tests for a product whose code you can read. FlowQA runs each test in a real browser against one of the project's environments, the way a person uses the product. A test you save is published at once: the person sees it in Tests and can run, edit, rename or delete it. Read this guide once per session; the quality checks on every save follow it.

1. What a test is

A test is one scenario a person would recognise, such as "pay for a cart with the saved card". It is a JSON Flow with ir: 1, the project's app id, platform: "web", a flow id and a name. The flow id is a path, module/slug, for example checkout/pay-with-saved-card: lowercase words joined by hyphens, no spaces. The name is what a person calls the test, written as a sentence. steps do what the person does, in order, and assertions say what the person sees. A test with an outline and no steps is planned: a request for a person to record it. A test marked reusable: true is a shared opening, such as signing in, that other tests include with a use step. rationale is one line on what the test proves; it is also where a choice the person asked for is recorded.

2. Find your way first

Call flowqa_get_project before writing anything. It gives you the app id, the environments and which of them are read-only, the modules, the suites, the shared openings and the variable names per environment. Put each test in a module the project already has; add a module only when nothing fits, with flowqa_organize and add_modules, or with create_modules: true on the save. When a shared opening already signs in or reaches the page you need, start with { "do": "use", "flow": "<its flow id>" } instead of writing those steps again. Call flowqa_list_tests with a module or a query to make sure the test you are about to write does not exist. When the listing answers truncated: true, the project has more tests than one listing returns, so narrow it with module, suite_id or query before you conclude a test is missing. If the test exists, read it with flowqa_get_test and improve it instead of adding a second one. flowqa_site_notes says what the project knows about each page; read the note for a page whose code you have not found.

3. Writing steps

Step kinds: navigate (a url, relative to the environment's base URL), reload, back, forward, tap, fill and select (a target and a value), press (a key, such as Enter), waitFor and assert (an assertion in that), snapshot (an extra screenshot, with a label) and use (a reusable flow). A target names an element the way the accessibility tree does: role and label together, for example { "role": "button", "label": "Pay now" }. Read the component in the repository to find both: the element's role, and its accessible name from the visible text of a button or link, the <label> of a field or an aria-label. Use text only when the element has no role and name. attrs (href, name) helps when the visible text changes between languages, but it goes alongside role and label or text, never alone: a target needs one of those or a fallback. When a label repeats, such as "Remove" on every cart row, scope it with within: containers outermost first, each with a role and a name or text, for example the row that holds the product's name. Use nth only as a last resort. Never make a CSS selector or a generated id the way a step finds its element. If nothing else exists, put a data-testid in fallback.testId; when you can change the product, add the accessible name or the test id there instead. A target found by text or fallback.testId alone always carries a weak_locator finding, and that finding stays until the product gives the element an accessible name. A fill into a password, one-time code or card field binds a variable, such as {{customer_password}}, never a literal.

4. Writing assertions

Every test with steps ends in at least one assertion about the outcome the person cares about. Assertions are objects with one key: textVisible and textHidden (a string), urlMatches (a regular expression, so escape . and ? when you mean them literally), elementVisible, elementHidden and elementDisabled (a target), elementValue ({ "target", "value" }), fieldError ({ "label", "message" }), toast (a string), networkStatus ({ "urlPattern", "code" }), noNewConsoleErrors (true) and backendCheck (a check the project defines). Prefer what the person sees: textVisible, urlMatches, elementVisible and toast. Use networkStatus only when the response itself is the point of the test. When the page takes a moment, give the assert a timeoutMs, or put a waitFor before the next action; the claim stays the same, the page just has longer to make it true. { "do": "assert", "that": { "noNewConsoleErrors": true } } is a cheap last step that catches a page that broke quietly.

5. Data and variables

Write {{customer_email}} in a url, a value or an assertion to use a variable. Variables are defined per environment in the project's settings; flowqa_get_project lists their names and never their values. A value only one test needs can go in that test's own data object. Never hard-code production data, a real customer's details or a credential. When a test needs a value you cannot find in the repository, such as a promo code that exists only on staging, ask the person to add the variable, or put the need in an outline and request a recording.

6. Environments

A test runs on every environment that is not read-only, unless it lists its own in environments. A read-only environment, usually production, runs only the tests that name it. Name production only for a test that changes nothing, and do not name only read-only environments unless the person asked for production checks. A step that belongs on some environments only, such as a preview gate, carries onlyOn with their ids.

7. Organising

Modules group tests by product area, such as auth, checkout or search. Suites group modules by purpose, such as a smoke suite or a release suite. order sets where a test sits in its module, in the sequence a person would work through it. Organise with flowqa_organize, passing expected_configuration_revision set to the configuration_revision that flowqa_get_project returns: move_tests moves tests into a module, rename_module renames one, set_order sets positions and suites replaces the list of suites. To rename a test, save it with a new flow id. A move or a rename keeps the old id in aka, so run history follows the test. Keep shared openings in a module of their own, such as auth, and mark them reusable: true. Delete tests with flowqa_delete_tests, passing test_ids and, in the same order, expected_revisions: the revision of each that you last read. Delete only what the person asked for or what you wrote; a shared opening that other tests still use is refused.

8. The loop

  1. Save with flowqa_save_tests; every test publishes at once as a revision and comes back with findings.
  2. Fix every finding you can and save again with test_id and base_revision set to the revision you got back. A weak_locator on an element with no accessible name stays until the product gets one; tell the person about it. check_only: true returns the findings without writing anything.
  3. Run with flowqa_run_tests, picking the tests with exactly one of test_ids, module, suite_id or all: true, and optionally a label that names the run. It runs on a cloud browser by default, or on the person's FlowQA Desktop with target: "local" when the environment is only reachable from their machine. You can also pass snag_id to run the tests linked to a snag, and environment to pick an environment other than the project's default. Planned tests are left out, and so, on a read-only environment, are tests that do not name it; the reply lists each as skipped, with the reason. Shared openings are never run on their own: one you pick by id is listed as skipped, and one a module, suite_id or all selection reaches is dropped silently.
  4. Poll flowqa_get_run with each execution_id until the run ends: completed, failed, cancelled or expired (a local run expires when no FlowQA Desktop picks it up).
  5. For a failure, read the failing step and fetch its screenshot with flowqa_get_artifact, using the path of the artifact of kind failing.
  6. Decide: the product is wrong (fix the code, or file it with file_as_snag), or the test is wrong (fix the step and save a new revision).
  7. Tell the person what you wrote, what passed, what you fixed and what you could not decide.

A save that conflicts (409) means someone changed the test since you read it: read it again with flowqa_get_test, merge your change into what is there, and save with base_revision set to its current_revision. Never work around a conflict by deleting and recreating the test. A batch with one invalid test saves nothing; the error names each invalid test and why, so fix them and send the batch again.

9. Working with the person

Some flows cannot be read from the code: a third-party payment widget, a captcha, a link in an email, a visual judgement. For those, call flowqa_request_recording with a flow_id, a name and an outline of what to do, one line per action, and set opening when the flow starts from a shared opening. Give the person the desktop_link it returns, which opens FlowQA Desktop on that planned test; web_link opens the test in the app and offers the download when the desktop is not installed. Poll flowqa_get_test until planned is false, then add the assertions the person did not record, and run it. Never use a secret the person has not given you, and never copy one from production configuration into a test. When an exploration asks you for a value, never give a secret the person has not shown you. When the value you give is a secret, mark it as secret and have it remembered as a project variable, so the next exploration does not ask again. When you cannot answer from the code or the screenshot, hand the question to the person. When you cannot tell whether a failure is the product or the test, ask the person before changing either.

10. What the person can override

Every rule in this guide except validity is a default the person can change. If the person tells you to keep a CSS selector, to skip an assertion or to run a test on production only, do it, and say so in the test's rationale so the next reader knows the finding is intended. Validity is not negotiable, and a save refuses a test that does not parse, has an app other than the project's, has a flow id another test already has, names a module or an environment the project does not have, or has a use step naming a test that does not exist, is not reusable or runs on another platform.

Example tests

A shared opening that signs in, marked reusable.

{
  "ir": 1,
  "app": "shop",
  "flow": "auth/sign-in",
  "name": "Sign in as the test customer",
  "platform": "web",
  "reusable": true,
  "rationale": "Every checkout test starts signed in as the staging test customer.",
  "tags": [],
  "setup": {},
  "data": {},
  "steps": [
    { "do": "navigate", "url": "/sign-in" },
    { "do": "fill", "target": { "role": "textbox", "label": "Email" }, "value": "{{customer_email}}" },
    { "do": "fill", "target": { "role": "textbox", "label": "Password" }, "value": "{{customer_password}}" },
    { "do": "tap", "target": { "role": "button", "label": "Sign in" } },
    { "do": "waitFor", "that": { "urlMatches": "/account" }, "timeoutMs": 10000 }
  ]
}

A test that starts from it and ends in what the customer sees.

{
  "ir": 1,
  "app": "shop",
  "flow": "checkout/pay-with-saved-card",
  "name": "Pay for a cart with the saved card",
  "platform": "web",
  "rationale": "A signed-in customer pays with the card on file and sees the order confirmation.",
  "order": 1,
  "tags": ["smoke"],
  "setup": {},
  "data": {},
  "steps": [
    { "do": "use", "flow": "auth/sign-in" },
    { "do": "navigate", "url": "/products/blue-mug" },
    { "do": "tap", "target": { "role": "button", "label": "Add to cart" } },
    { "do": "assert", "that": { "toast": "Added to cart" } },
    { "do": "tap", "target": { "role": "link", "label": "Cart" } },
    { "do": "tap", "target": { "role": "button", "label": "Checkout", "within": [{ "role": "region", "name": "Order summary" }] } },
    { "do": "tap", "target": { "role": "radio", "label": "Saved card ending 4242" } },
    { "do": "tap", "target": { "role": "button", "label": "Pay now" } },
    { "do": "assert", "that": { "urlMatches": "/orders/" }, "timeoutMs": 15000 },
    { "do": "assert", "that": { "textVisible": "Thank you for your order" } },
    { "do": "assert", "that": { "noNewConsoleErrors": true } }
  ]
}

A planned test, waiting for a person to record it in FlowQA Desktop.

{
  "ir": 1,
  "app": "shop",
  "flow": "checkout/apply-promo-code",
  "name": "Apply a promo code at checkout",
  "platform": "web",
  "rationale": "The total drops by the promo amount and the code shows as applied.",
  "outline": [
    "Start from the shared opening \"auth/sign-in\".",
    "Add the blue mug to the cart.",
    "Open the cart and type the promo code from the promo_code variable.",
    "Apply it and check the new total."
  ],
  "tags": [],
  "setup": {},
  "data": {},
  "steps": []
}

When an agent reads this guide through the flowqa_guide tool, the IR JSON Schema of a test follows here, generated from the test format the API checks.