TestFinch

Docs

The explore-click benchmark

What the 97.5% on the FlowQA page measures, on which sites, how right is counted, what it does not show, and how to run it.

FlowQA's exploring agent takes one decision at a time: given a goal and the interactive elements on the page, which element does it act on next? A strong model answers that well and slowly, at a price that adds up over a few hundred decisions per site. This benchmark asked whether a cheap decision model could answer as well, with the strong model kept for the decisions it is unsure about. The product page quotes one number from it, 97.5%. This document says exactly what that number measures, what it does not, and how to run the benchmark yourself.

The question

During automated exploration of a website, can a cheap decision model (Jev 1.13 by TypeSafe AI) choose the right next UI element as well as a small general model (Claude Haiku 4.5), with a strong model (Claude Sonnet 5) as the reference ceiling?

The dataset

88 journeys across ten sites, replayed end to end in a headless browser on 2026-10-07.

Site Journeys Element steps
A client's car-parts shop (development site, behind a preview gate; aggregate rows only) 18 36
saucedemo 10 28
FlowQA's own demo shop 10 19
the-internet (Heroku test app) 13 19
TodoMVC 6 17
demoqa 8 12
automationexercise 6 11
books.toscrape 6 8
quotes.toscrape 5 7
playwright.dev docs 6 7

Journeys are written the way FlowQA's planner writes them: a goal in one sentence, then the steps a person would take. Two of the client-site journeys are FlowQA's own recorded flows (the preview gate and a signed-in session); the rest were authored read-only against that development site, never placing an order or changing account data. The public sites were authored and validated headless. Every step's locator was checked to resolve to exactly one visible element before any model was called.

What a decision point is

Before each click, fill or select, the harness records what the model sees and what the right answer is.

That gives 164 element steps. One more decision point is added at the end of each journey, where the right answer to "is the goal met" is yes, for 252 decision points in all. Secrets filled during a journey never reach a model; they appear as <secret:name> in every state and are scrubbed from candidate text.

How right is counted

A choice is right when it is the recorded element, another link with the same destination, or the one other element on the page with the identical role and name (a card's image button and its title button). A name shared by three or more elements is not treated as interchangeable. The "all steps" column counts an element the extractor failed to list as a miss; the "truth in candidates" column drops those steps from the denominator. One step of 164 had its element beyond the 150-candidate cap, so the two columns differ by one step.

Results

Variant Model Top-1, all 164 steps Top-1, 163 with truth listed Goal-met accuracy Median latency Cost per 1,000 decisions
jev-flat Jev 1.13 97.0% 97.5% 92.1% 350 ms $0.23
haiku Claude Haiku 4.5 97.0% 97.5% 89.7% 1,643 ms $3.03
sonnet Claude Sonnet 5 97.0% 97.5% 92.9% 1,404 ms $6.76

Three hierarchical variants (choose a region of the page, then an element inside it) scored between 92.0% and 95.1% and are in the full results. Cost is the API-reported token usage at list prices on the day.

Each of the three models missed four of the 163 steps, not the same four. The cheap model missed two steps on demoqa and two on the client site. Sonnet missed one on demoqa, two on the client site and one on TodoMVC. Haiku missed one on demoqa, one on saucedemo and two on the client site. Every miss on the public sites, with the state the model saw and the answer it gave, is in the published answer files named below; the client site's misses are described only in aggregate.

What this does and does not show

It shows that, on these 163 decisions, the cheap decision model picks the recorded element as often as the strong model, about four times faster and at a thirtieth of the cost. That is the basis for FlowQA using it for every routine decision and escalating to the strong model when its confidence is low.

It does not show that whole journeys succeed at that rate. A journey of eight steps at 97.5% per step is right end to end about 82% of the time if the misses were independent, which is why the explorer verifies every generated test on fresh sessions and asks a person when it is unsure. It does not cover pages that need hovering, drag, keyboard-only interaction or file upload, which the explorer does not do. It is one run of 164 steps on ten sites, eight of them public demo sites that are simpler than most production applications. We will add sites and re-run it; this page will carry the date of the latest run.

The published package

Everything a reader needs to check the numbers is published beside this page, with the client site's rows removed and its aggregate rows kept under the label client-site.

A decision point's id joins the answer files to the decision points, so any row in the tables can be traced to the state the model saw and the answer it gave.

What is not published

The harness itself (the extractor, the replay, the model backends) lives in TestFinch's private repository. This page is therefore a published methodology with published data, not a package a stranger can re-run end to end. If you want to re-run it, write to hello@theproducthighway.com and we will share the harness under a short agreement. The client site's decision points and answers are withheld because they contain that site's page text.