TestFinch

Blog

Visual regression testing without the noise

2026-10-08 ยท Screenshot comparison finds the breaks nobody wrote an assertion for, and drowns them in font rendering and animation. How to keep the signal.

Visual regression testing works when three things hold: baselines come from the same rendering environment as the runs, visual drift is reported separately from functional failure, and a person approves a new baseline from a run they have looked at. Without those three, the comparison reports every font hint and every half-finished animation, and the team stops reading it within a week.

What visual checks catch that assertions do not

A functional test asserts what it was told to. It does not notice that the sidebar now overlaps the content, that the dark theme lost its contrast, or that a button is rendered off-screen but still "visible" to the test. A screenshot of each step, compared against a known-good one, sees all of that. That is the signal.

Where the noise comes from

Rendering differences. The same page on a Mac and on a Linux worker renders fonts differently. Line breaks move, buttons shift a pixel, the diff lights up. Comparing a baseline from one environment against a run from another produces noise on every run.

Motion. A screenshot taken mid-transition differs from one taken at rest. Spinners, fades, carousels.

Dynamic content. Dates, counters, avatars, anything that changes between runs.

Everything at once. A full-page diff of a dashboard with ten widgets will differ somewhere on most runs.

Keeping the signal

Compare like with like. A baseline is valid for one rendering profile: operating system, browser build, viewport, and the fonts installed. When the worker image changes, the profile changes, and the baselines are re-approved, deliberately. Never compare a cloud run against a screenshot from a laptop.

Settle before shooting. Wait for the network to go quiet and animations to finish. A tool that captures on a fixed delay will catch motion sooner or later.

Report drift as drift. A run where every assertion passed and three screenshots differ is a different thing from a run where a step failed. Show it as such. A failing test is a stop; visual drift is a review. Mixing them in one red mark is what makes people ignore both.

Approve from a run you looked at. A new baseline should come from a completed run on the current version of the tests, after a person has seen the differences and agreed they are intended. A run with functional failures cannot be a baseline; a run whose only differences are drift can, once reviewed.

Mask what moves. Dates and counters are masked or asserted separately.

A worked example

A composed example: the team, the timings and the counts show the shape of the work and are not measurements from one customer.

A team adds visual checks to the twelve journeys in their merge gate. The first week is loud: every run reports drift on every page, because the baselines were captured on a developer's Mac and the runs are on Linux workers. They re-approve the baselines from a cloud run, and the noise stops.

Two weeks later a run passes every assertion and reports drift on the settings page. The diff shows the save button has moved under the footer at the tablet viewport after a CSS change. No assertion would have caught it; the button was still clickable to the test. They fix the CSS, the next run shows no drift, and nobody had to approve anything.

FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.

See FlowQA