TestFinch

Blog

Playwright vs record-and-replay: when each wins

2026-10-08 ยท Hand-written Playwright and recorded tests are not rivals. One is for the code you own in depth, the other for the product everyone touches. Here is where the line falls.

Write Playwright by hand for the parts of your product that have complex state and a developer who owns them. Record for everything else: the journeys that many people change and nobody owns, where the cost of writing and maintaining code is what stops the tests from existing at all. Most teams need both, and the mistake is choosing one as a policy.

What hand-written Playwright is good at

Playwright gives you a programming language. That matters when a test needs a loop, a fixture that seeds data through an API, a mock of a third-party service, or an assertion that compares two things the page never shows side by side. It is version-controlled next to the code, runs in the same pipeline as unit tests, and a developer can debug it with a breakpoint.

It is also work. Every selector is a decision. Every page object is code someone maintains. A team of three developers can keep fifty such tests healthy. A team of three cannot keep five hundred, and the tests that are not healthy get skipped, then deleted.

What record-and-replay is good at

Recording captures what a person did: the clicks, the inputs, the pages, and the element each action landed on. It produces a test in minutes from someone who knows the product rather than someone who knows the framework. A product manager can record the journey they care about. A support engineer can record the bug they just reproduced.

Early recorders earned a bad name because they recorded brittle selectors and no assertions. Modern ones record the element by several handles at once (role, text, test id, position in its form) and pick the one that still works at replay time. They prompt for assertions, or infer them from what changed on the page. And the better ones go further: an agent explores the site, proposes the journeys, writes the tests, and proves each one by breaking the page and watching the test fail.

Where the line falls

Use Playwright when the test needs code: data setup through an API, mocked payments, parallel users, time travel. Use recording when the test is a journey a person could walk: sign up, create, search, check out, change a setting. In most products the second group is ten times larger than the first.

Keep hand-written tests for the core state machine and the integration edges. Let recorded and generated tests cover the surface, where change is constant and ownership is shared.

A worked example

A composed example: the team, the timings and the counts show the shape of the work and are not measurements from one customer.

An e-commerce team has a cart with promotions, a known source of regressions. Their Playwright suite covers promotion rules: twelve tests that seed products through the admin API, apply codes and assert totals to the cent. Those stay hand-written; the arithmetic needs code.

The same team had no tests for the account pages, because no developer owned them and nobody wanted to write page objects for a settings form. They recorded nine journeys in an afternoon: change email, add an address, set a default card, export data. When an agent later reworked the address form, the "add an address" test failed, the tool proposed a one-line repair for the renamed field, and a person approved it. Twelve tests in code, nine recorded, both in the pipeline, each where it belongs.

The one rule

Whichever way a test is made, it must be able to fail. Run it against a broken page once before you trust a green from it.

FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.

See FlowQA