The explore-click benchmark
FlowQA's exploring agent takes one decision at a time: given a goal and the interactive elements on the page, which element does it act on next? A strong model answers that well and slowly, at a price that adds up over a few hundred decisions per site. This benchmark asked whether a cheap decision model could answer as well, with the strong model kept for the decisions it is unsure about. The product page quotes one number from it, 97.5%. This document says exactly what that number measures, what it does not, and how to run the benchmark yourself.
The question
During automated exploration of a website, can a cheap decision model (Jev 1.13 by TypeSafe AI) choose the right next UI element as well as a small general model (Claude Haiku 4.5), with a strong model (Claude Sonnet 5) as the reference ceiling?
The dataset
88 journeys across ten sites, replayed end to end in a headless browser on 2026-10-07.
| Site | Journeys | Element steps |
|---|---|---|
| A client's car-parts shop (development site, behind a preview gate; aggregate rows only) | 18 | 36 |
| saucedemo | 10 | 28 |
| FlowQA's own demo shop | 10 | 19 |
| the-internet (Heroku test app) | 13 | 19 |
| TodoMVC | 6 | 17 |
| demoqa | 8 | 12 |
| automationexercise | 6 | 11 |
| books.toscrape | 6 | 8 |
| quotes.toscrape | 5 | 7 |
| playwright.dev docs | 6 | 7 |
Journeys are written the way FlowQA's planner writes them: a goal in one sentence, then the steps a person would take. Two of the client-site journeys are FlowQA's own recorded flows (the preview gate and a signed-in session); the rest were authored read-only against that development site, never placing an order or changing account data. The public sites were authored and validated headless. Every step's locator was checked to resolve to exactly one visible element before any model was called.
What a decision point is
Before each click, fill or select, the harness records what the model sees and what the right answer is.
- Candidates: the visible, enabled interactive elements on the page, each with a key, a role, an accessible name, the text of the row or card it sits in, the nearest heading, its landmark, and whether it is in the viewport. Viewport-first order, capped at 150 per page.
- Truth: the element the recorded step acts on, matched by DOM identity.
- The state the model gets: the goal, the page title and path, the history so far ("step 1: clicked link Products"), and one line per candidate. Median state is about 500 tokens; the largest is under 7,000.
That gives 164 element steps.
One more decision point is added at the end of each journey, where the right answer to "is the goal met" is yes, for 252 decision points in all.
Secrets filled during a journey never reach a model; they appear as <secret:name> in every state and are scrubbed from candidate text.
How right is counted
A choice is right when it is the recorded element, another link with the same destination, or the one other element on the page with the identical role and name (a card's image button and its title button). A name shared by three or more elements is not treated as interchangeable. The "all steps" column counts an element the extractor failed to list as a miss; the "truth in candidates" column drops those steps from the denominator. One step of 164 had its element beyond the 150-candidate cap, so the two columns differ by one step.
Results
| Variant | Model | Top-1, all 164 steps | Top-1, 163 with truth listed | Goal-met accuracy | Median latency | Cost per 1,000 decisions |
|---|---|---|---|---|---|---|
| jev-flat | Jev 1.13 | 97.0% | 97.5% | 92.1% | 350 ms | $0.23 |
| haiku | Claude Haiku 4.5 | 97.0% | 97.5% | 89.7% | 1,643 ms | $3.03 |
| sonnet | Claude Sonnet 5 | 97.0% | 97.5% | 92.9% | 1,404 ms | $6.76 |
Three hierarchical variants (choose a region of the page, then an element inside it) scored between 92.0% and 95.1% and are in the full results. Cost is the API-reported token usage at list prices on the day.
Each of the three models missed four of the 163 steps, not the same four. The cheap model missed two steps on demoqa and two on the client site. Sonnet missed one on demoqa, two on the client site and one on TodoMVC. Haiku missed one on demoqa, one on saucedemo and two on the client site. Every miss on the public sites, with the state the model saw and the answer it gave, is in the published answer files named below; the client site's misses are described only in aggregate.
What this does and does not show
It shows that, on these 163 decisions, the cheap decision model picks the recorded element as often as the strong model, about four times faster and at a thirtieth of the cost. That is the basis for FlowQA using it for every routine decision and escalating to the strong model when its confidence is low.
It does not show that whole journeys succeed at that rate. A journey of eight steps at 97.5% per step is right end to end about 82% of the time if the misses were independent, which is why the explorer verifies every generated test on fresh sessions and asks a person when it is unsure. It does not cover pages that need hovering, drag, keyboard-only interaction or file upload, which the explorer does not do. It is one run of 164 steps on ten sites, eight of them public demo sites that are simpler than most production applications. We will add sites and re-run it; this page will carry the date of the latest run.
The published package
Everything a reader needs to check the numbers is published beside this page, with the client site's rows removed and its aggregate rows kept under the label client-site.
- summary.md: the full results, with calibration bins and per-site breakdowns.
- metrics.json: every figure in the tables, as data.
- backend-info.json: the exact model ids and endpoints each variant used.
- decision-points.jsonl: the 198 public-site decision points, each with the goal, the page, the history and the candidate list the models saw, and the truth.
- Answers, one file per variant, one line per decision point: jev-flat, jev-hier, jev-fanout, haiku, haiku-hier, sonnet.
- The nine public-site journey files under /benchmark/dataset/ (demo-app, saucedemo, the-internet, todomvc, demoqa, automationexercise, toscrape, quotes, playwright-docs), in the harness's dataset format.
A decision point's id joins the answer files to the decision points, so any row in the tables can be traced to the state the model saw and the answer it gave.
What is not published
The harness itself (the extractor, the replay, the model backends) lives in TestFinch's private repository. This page is therefore a published methodology with published data, not a package a stranger can re-run end to end. If you want to re-run it, write to hello@theproducthighway.com and we will share the harness under a short agreement. The client site's decision points and answers are withheld because they contain that site's page text.