# Explore-click benchmark: results

Generated 2026-10-07T14:58:32.637Z.
Question: during automated exploration, can Jev pick the next UI element as well as Haiku 4.5, with Sonnet 5 as the ceiling?
Variants: `jev-flat` (one choice over every candidate), `jev-hier` (two serial requests: region, then element inside it), `jev-fanout` (one request: region choice plus one best-element question per region), `haiku`, `haiku-hier` (two serial calls), `sonnet` (flat, the ceiling).

## Dataset and extractor

- Journeys: 88; decision points: 252 (164 element steps + 88 end-of-journey goal checks).
- Truth in candidates: 163 of 164 steps (99.4%).
- Truth not in candidates: 0 (0.0%); truth cut off by the 150-candidate cap: 1 (0.6%).
- Truth matched by: self 159, ancestor 3, label-for 2.
- Candidates per page (sent): median 22, p95 150, max 150; before the cap: median 22, p95 180, max 388.
- State size (approx tokens, flat prompt): median 523, p95 6456, max 6820.
- Regions per page: median 4, p95 12, max 12; region kinds: section 594, header 178, nav 148, footer 139, form 116, page 72, other 47, main 31, aside 28, dialog 5, overlay 4.
- Truth element's region found (a named region, not "other"): 98.2%; truth region kinds: section 51, form 41, header 22, nav 17, main 11, footer 5, aside 5, page 5, other 3, dialog 2, overlay 1; elements in the truth's region: median 4, p95 44.

| Site | Journeys | Decision points | Element steps | Truth in candidates | Not in candidates | Cut off |
|---|---:|---:|---:|---:|---:|---:|
| automationexercise | 6 | 17 | 11 | 11 | 0 | 0 |
| books-toscrape | 6 | 14 | 8 | 8 | 0 | 0 |
| demo-app | 10 | 29 | 19 | 19 | 0 | 0 |
| demoqa | 8 | 20 | 12 | 12 | 0 | 0 |
| playwright-docs | 6 | 13 | 7 | 6 | 0 | 1 |
| quotes-toscrape | 5 | 12 | 7 | 7 | 0 | 0 |
| saucedemo | 10 | 38 | 28 | 28 | 0 | 0 |
| client-site | 18 | 54 | 36 | 36 | 0 | 0 |
| the-internet | 13 | 32 | 19 | 19 | 0 | 0 |
| todomvc | 6 | 23 | 17 | 17 | 0 | 0 |

## Headline

| Variant | Model | Top-1 (all steps) | Top-1 (truth in candidates) | Top-3 | Action type | Goal-met acc | Goal-met FPR | ECE | Round trips | Latency p50 / p95 | Cost per 1,000 decisions | Errors |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| jev-flat | jev-1.13.0 | 97.0% | 97.5% | 100.0% | 98.2% | 92.1% | 0.0% | 0.019 | 1 | 350 / 524 ms | $0.23 | 0 |
| jev-hier | jev-1.13.0 | 93.9% | 94.5% | 95.8% | 96.3% | 91.3% | 0.0% | 0.039 | 2 | 642 / 863 ms | $0.12 | 0 |
| jev-fanout | jev-1.13.0 | 94.5% | 95.1% | 96.6% | 97.6% | 91.3% | 0.0% | 0.006 | 1 | 359 / 548 ms | $0.28 | 0 |
| haiku | claude-haiku-4-5-20251001 | 97.0% | 97.5% | n/a | 96.3% | 89.7% | 0.0% | 0.016 | 1 | 1643 / 3287 ms | $3.03 | 0 |
| haiku-hier | claude-haiku-4-5-20251001 | 91.5% | 92.0% | n/a | 96.3% | 81.7% | 3.7% | 0.033 | 2 | 2926 / 5031 ms | $2.17 | 0 |
| sonnet | claude-sonnet-5 | 97.0% | 97.5% | n/a | 97.6% | 92.9% | 0.0% | 0.033 | 1 | 1404 / 1825 ms | $6.76 | 0 |

Denominators: 164 element steps, 163 with truth in candidates, 252 goal checks (one per decision point).
Top-3 uses the returned probability distribution and is only available for backends that return one (Jev); the Claude backends return a single choice.
A choice counts as right when it is the recorded element, another link with the same href, or the one other element with the identical role and name (a card image button and its title button).
Claude answers that came back as prose instead of JSON were asked once more for the object alone (counted in cost and latency): haiku 1, haiku-hier 2.
Latency is wall-clock per decision from this machine; two-round-trip variants add their serial calls.

## Hierarchical variants

| Variant | Region accuracy | Element accuracy given region right | End-to-end top-1 | Alternative rule (max region x element prob) | Latency parts p50 (ms) | Request size p50 / max (bytes) | Input tokens p50 / p95 |
|---|---:|---:|---:|---:|---|---|---|
| jev-hier | 94.5% | 98.1% | 94.5% | n/a | 309 + 316 | 4838 / 37513 | 1943 / 6313 |
| jev-fanout | 95.1% | 97.4% | 95.1% | 95.1% | single | 8322 / 70397 | 2868 / 20704 |
| haiku-hier | 92.6% | 97.4% | 92.0% | n/a | 1675 + 1121 | n/a | 878 / 2994 |

Region accuracy is over the 163 steps with truth in candidates. The Jev context limit is 32K tokens for state plus the longest question; request sizes above are what was sent.

## Top-1 per site (truth in candidates)

| Site | Steps | jev-flat | jev-hier | jev-fanout | haiku | haiku-hier | sonnet |
|---|---:|---:|---:|---:|---:|---:|---:|
| automationexercise | 11 | 100.0% | 90.9% | 100.0% | 100.0% | 90.9% | 100.0% |
| books-toscrape | 8 | 100.0% | 87.5% | 100.0% | 100.0% | 87.5% | 100.0% |
| demo-app | 19 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% |
| demoqa | 12 | 83.3% | 75.0% | 83.3% | 91.7% | 83.3% | 91.7% |
| playwright-docs | 6 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% |
| quotes-toscrape | 7 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% |
| saucedemo | 28 | 100.0% | 100.0% | 100.0% | 96.4% | 96.4% | 100.0% |
| client-site | 36 | 94.4% | 88.9% | 86.1% | 94.4% | 86.1% | 94.4% |
| the-internet | 19 | 100.0% | 100.0% | 94.7% | 100.0% | 100.0% | 100.0% |
| todomvc | 17 | 100.0% | 100.0% | 100.0% | 100.0% | 82.4% | 94.1% |

## Calibration of the chosen option

Accuracy per bin of the probability the backend assigned to its own choice (Jev flat: the returned probability; hierarchical: region probability x element probability; Claude: the self-reported confidence number).

### jev-flat (ECE 0.019, n=163)

| Bin | n | Accuracy | Avg confidence |
|---|---:|---:|---:|
| 0.0-0.2 | 0 | n/a | n/a |
| 0.2-0.4 | 0 | n/a | n/a |
| 0.4-0.6 | 4 | 75.0% | 0.55 |
| 0.6-0.8 | 7 | 85.7% | 0.69 |
| 0.8-1.0 | 152 | 98.7% | 0.98 |

### jev-hier (ECE 0.039, n=163)

| Bin | n | Accuracy | Avg confidence |
|---|---:|---:|---:|
| 0.0-0.2 | 0 | n/a | n/a |
| 0.2-0.4 | 1 | 100.0% | 0.38 |
| 0.4-0.6 | 5 | 60.0% | 0.50 |
| 0.6-0.8 | 12 | 91.7% | 0.69 |
| 0.8-1.0 | 145 | 95.9% | 0.98 |

### jev-fanout (ECE 0.006, n=163)

| Bin | n | Accuracy | Avg confidence |
|---|---:|---:|---:|
| 0.0-0.2 | 0 | n/a | n/a |
| 0.2-0.4 | 0 | n/a | n/a |
| 0.4-0.6 | 2 | 50.0% | 0.54 |
| 0.6-0.8 | 9 | 66.7% | 0.73 |
| 0.8-1.0 | 152 | 97.4% | 0.97 |

### haiku (ECE 0.016, n=163)

| Bin | n | Accuracy | Avg confidence |
|---|---:|---:|---:|
| 0.0-0.2 | 0 | n/a | n/a |
| 0.2-0.4 | 0 | n/a | n/a |
| 0.4-0.6 | 0 | n/a | n/a |
| 0.6-0.8 | 0 | n/a | n/a |
| 0.8-1.0 | 163 | 97.5% | 0.96 |

### haiku-hier (ECE 0.033, n=163)

| Bin | n | Accuracy | Avg confidence |
|---|---:|---:|---:|
| 0.0-0.2 | 2 | 0.0% | 0.12 |
| 0.2-0.4 | 0 | n/a | n/a |
| 0.4-0.6 | 0 | n/a | n/a |
| 0.6-0.8 | 1 | 100.0% | 0.71 |
| 0.8-1.0 | 160 | 93.1% | 0.90 |

### sonnet (ECE 0.033, n=163)

| Bin | n | Accuracy | Avg confidence |
|---|---:|---:|---:|
| 0.0-0.2 | 0 | n/a | n/a |
| 0.2-0.4 | 0 | n/a | n/a |
| 0.4-0.6 | 0 | n/a | n/a |
| 0.6-0.8 | 3 | 66.7% | 0.70 |
| 0.8-1.0 | 160 | 98.1% | 0.95 |

## Escalation policy

If the cheap variant answers alone when its confidence is at least T and otherwise defers to Sonnet: accuracy and share of steps escalated, over steps with truth in candidates.

| T | jev-flat acc | jev-flat esc | jev-hier acc | jev-hier esc | jev-fanout acc | jev-fanout esc | haiku acc | haiku esc | haiku-hier acc | haiku-hier esc |
|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| 0.50 | 97.5% | 0.6% | 95.1% | 2.5% | 95.1% | 0.6% | 97.5% | 0.0% | 93.3% | 1.2% |
| 0.55 | 97.5% | 0.6% | 95.7% | 3.1% | 95.1% | 0.6% | 97.5% | 0.0% | 93.3% | 1.2% |
| 0.60 | 98.2% | 2.5% | 95.7% | 3.7% | 95.7% | 1.2% | 97.5% | 0.0% | 93.3% | 1.2% |
| 0.65 | 98.2% | 3.1% | 95.7% | 4.9% | 95.7% | 2.5% | 97.5% | 0.0% | 93.3% | 1.2% |
| 0.70 | 98.2% | 4.9% | 95.1% | 8.0% | 95.7% | 3.1% | 97.5% | 0.0% | 93.3% | 1.2% |
| 0.75 | 98.2% | 6.1% | 95.1% | 9.8% | 96.3% | 3.7% | 97.5% | 0.0% | 93.3% | 1.8% |
| 0.80 | 98.2% | 6.7% | 95.7% | 11.0% | 97.5% | 6.7% | 97.5% | 0.0% | 93.3% | 1.8% |
| 0.85 | 98.2% | 9.2% | 95.7% | 13.5% | 97.5% | 11.7% | 97.5% | 0.0% | 94.5% | 5.5% |
| 0.90 | 98.2% | 13.5% | 96.9% | 16.0% | 98.2% | 16.0% | 98.2% | 1.2% | 94.5% | 5.5% |
| 0.95 | 97.5% | 19.0% | 97.5% | 25.2% | 98.2% | 23.3% | 98.2% | 1.2% | 97.5% | 100.0% |

Sonnet alone: 97.5% at 100% escalation.

## Example misses

### jev-flat

- demoqa /checkbox (step 1) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c4 clickable ".rc-tree-switcher"; chosen: c3 clickable "Home" (p=0.59) [truth region r3, chosen region r3]
- demoqa /checkbox (step 2) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c7 clickable "Desktop"; chosen: c9 checkbox "Select Desktop" (p=0.85) [truth region r3, chosen region r3]

### jev-hier

- books-toscrape /catalogue/category/books/poetry_23/index.html (step 2) - goal: "Open the Poetry category and open the second book listed."
  truth: c29 link "The Black Maria" (row: £52.15 In stock Add to basket); chosen: c28 link "A Light in the ..." (row: £51.77 In stock Add to basket) (p=0.50) [truth region r4, chosen region r3]
- demoqa /checkbox (step 1) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c4 clickable ".rc-tree-switcher"; chosen: c3 clickable "Home" (p=0.90) [truth region r3, chosen region r3]
- demoqa /checkbox (step 2) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c7 clickable "Desktop"; chosen: c9 checkbox "Select Desktop" (p=0.84) [truth region r3, chosen region r3]

### jev-fanout

- the-internet /dynamic_controls (step 1) - goal: "On the dynamic controls page, remove the checkbox and then enable the text input."
  truth: c2 button "Remove" (row: Dynamic Controls This example demonstrates when elements (e…); chosen: c1 checkbox "" (row: Dynamic Controls This example demonstrates when elements (e…) (p=0.89) [truth region r1, chosen region r1]
- demoqa /checkbox (step 1) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c4 clickable ".rc-tree-switcher"; chosen: c3 clickable "Home" (p=0.76) [truth region r3, chosen region r3]
- demoqa /checkbox (step 2) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c7 clickable "Desktop"; chosen: c9 checkbox "Select Desktop" (p=0.80) [truth region r3, chosen region r3]

### haiku

- saucedemo /inventory.html (step 5) - goal: "Add both the Sauce Labs Fleece Jacket and the Sauce Labs Onesie to the cart and open the cart."
  truth: c20 button "Add to cart" (row: Sauce Labs Onesie Rib snap infant onesie for the junior aut…); chosen: c15 button "Remove" (row: Sauce Labs Fleece Jacket It's not every day that you come a…) (p=0.85) [truth region r2, chosen region r2]
- demoqa /checkbox (step 2) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c7 clickable "Desktop"; chosen: c9 checkbox "Select Desktop" (p=0.95) [truth region r3, chosen region r3]

### haiku-hier

- todomvc /todomvc/#/ (step 3) - goal: "Add the todo 'Buy milk' and mark it as completed."
  truth: c4 checkbox "Toggle Todo" (row: Buy milk); chosen: c2 textbox "What needs to be done?" (p=0.90) [truth region r3, chosen region r2]
- todomvc /todomvc/#/ (step 3) - goal: "Add 'Buy milk', mark it completed, then clear the completed todos."
  truth: c4 checkbox "Toggle Todo" (row: Buy milk); chosen: c2 textbox "What needs to be done?" (p=0.90) [truth region r3, chosen region r2]
- todomvc /todomvc/#/ (step 3) - goal: "Add 'Buy milk', mark it completed, then open the Completed filter."
  truth: c4 checkbox "Toggle Todo" (row: Buy milk); chosen: c2 textbox "What needs to be done?" (p=0.90) [truth region r3, chosen region r2]
- saucedemo /inventory.html (step 5) - goal: "Add both the Sauce Labs Fleece Jacket and the Sauce Labs Onesie to the cart and open the cart."
  truth: c20 button "Add to cart" (row: Sauce Labs Onesie Rib snap infant onesie for the junior aut…); chosen: c15 button "Remove" (row: Sauce Labs Fleece Jacket It's not every day that you come a…) (p=0.81) [truth region r2, chosen region r2]
- books-toscrape /catalogue/category/books/poetry_23/index.html (step 2) - goal: "Open the Poetry category and open the second book listed."
  truth: c29 link "The Black Maria" (row: £52.15 In stock Add to basket); chosen: c16 link "Shakespeare's Sonnets" (row: Shakespeare's Sonnets £20.66 In stock Add to basket) (p=0.90) [truth region r4, chosen region r4]

### sonnet

- todomvc /todomvc/#/ (step 3) - goal: "Add 'Buy milk', mark it completed, then clear the completed todos."
  truth: c4 checkbox "Toggle Todo" (row: Buy milk); chosen: c2 textbox "What needs to be done?" (p=0.60) [truth region r3, chosen region r2]
- demoqa /checkbox (step 2) - goal: "Expand the Home node of the tree and tick the 'Desktop' checkbox."
  truth: c7 clickable "Desktop"; chosen: c9 checkbox "Select Desktop" (p=0.95) [truth region r3, chosen region r3]

## What this does and does not show

- It measures one narrow skill: given a goal, the history and a flat list of interactive elements, name the element a human tester used next. It does not measure planning over several steps, recovery from a wrong click, reading page content that is not an interactive element, or visual understanding.
- Ground truth is the single element the journey author used (plus links with the same href). Pages often offer several equally valid elements (a header and a footer link, a card button and its title). Those count as misses for every backend, so absolute numbers understate all of them; the comparison between variants is the useful part.
- Candidates are text-only descriptions from a cheap DOM extractor, capped at 150 viewport-first. Elements the extractor misses or cuts off are reported separately and count as misses in the "all steps" column; a better extractor lifts every backend.
- Regions come from the same extractor (dialog, landmarks, heading sections). A wrong region makes the hierarchical variants miss no matter how good the element choice is, so their numbers bound what segmentation quality allows.
- Journeys are short, read-only tasks on demo sites and one real car-parts shop (a client's development site behind a preview gate; only its aggregate rows are published here, labelled client-site). They are not a random sample of the web, and the public demo sites may be in LLM training data.
- Costs use list prices on the day of the run (Sonnet 5: $2 in / $10 out per million tokens; Jev 1.13: $0.042 in per million, output free) and the token counts the APIs reported; latency is wall-clock from this machine with concurrency 2 and includes network time.
- The goal-met question is heavily imbalanced (true only on the last page of each journey), so read accuracy together with the false-positive rate.
- Jev was reached via: typesafe https://api.typesafe.ai/v1/systemone model=jev-1.13.0.
