# a green test suite never opened the page

2026-09-30, updated 2026-09-30

By Jawad Jalal. Founder of Wayari. Builds the desktop app and its coding-agent workflow.

passing tests prove the cases they exercise. a browser check follows the person through the feature, including the part where they try again.


Passing tests show that the tested cases passed in the environment where they ran. A browser check shows whether a person can complete the changed task in the running application.

Both are useful. They answer different questions.

A login handler can return the right error while the page renders it behind a dialog. A button can call the right function while another element covers it on a phone.

Neither defect needs a broken unit test.

## what should you do in the browser first?

For a login fix, open the login page. Enter a wrong password. Read the error. Enter a correct one and try again.

For CSV export, open the report, select the rows, press export and inspect the resulting file. Counting a button click leaves the actual export untested.

Write that route through the feature into [the brief](/blog/how-to-write-a-brief-for-a-coding-agent). The builder and the verifier should be checking the same job.

## which environment and starting state should you use?

A preview is useful because it puts the branch in a running environment. Check that it contains the head you intend to review.

Then check the starting state. A signed-in developer with cached data may never meet the screen a new visitor sees.

Try a clean session for a first-run change. For a returning-user change, reproduce the saved state it depends on. A database full of convenient fixtures can hide a missing empty state.

Say which environment you used in the evidence. A local check that passed with development credentials leaves the production configuration as a separate question.

## how do you test recovery after a failure?

A useful check follows the person after something fails.

Does the spinner stop? Is the error readable? Can the person retry without refreshing? Does a second click submit twice?

For a network feature, simulate a refused or interrupted request where your tools allow it. For an input feature, use the values that caused the original defect.

The recovery path is part of the behaviour. A page that explains the error and leaves the person stuck still has a bug.

## what should you check on a phone and with a keyboard?

Open the changed view at a phone width if people use it on phones. Use the keyboard if the view has controls.

A screenshot can show clipping. It cannot tell you that focus disappears after the dialog closes or that a covered button will not accept a click.

Keep visual evidence and interaction evidence together. Record what you did and what the application returned, with enough detail that another person can repeat it.

## does a browser pass replace automated checks?

Manual use does not replace typecheck, builds or automated tests. A browser pass through one flow can miss a regression elsewhere.

Wayari's [gate](/blog/loop-until-it-holds) records the machine checks. Its workflow also asks agents to run the app and exercise the requested behaviour. The [reviewer](/blog/the-reviewer-is-never-the-author) then reads the change independently.

If a check could not run, leave that visible. "The browser was unavailable" is useful evidence about a limit. "Verified" hides the limit and leaves the next reader guessing.

## what does each kind of check establish?

Keep the evidence close to the claim it supports.

| Check | Evidence it provides | Example it can miss |
| --- | --- | --- |
| Unit test | A function handles the tested input. | The page hides its error message. |
| Browser interaction | The exercised flow works in that environment. | A regression in an unvisited flow. |
| Screenshot | The captured layout looked this way at that size. | Focus disappears after closing a dialog. |
| Production smoke check | The deployed route and configuration answer the exercised request. | An untested account or data state. |

## when should a browser check become a test?

When a failure is likely to recur and the behaviour can be asserted reliably, keep the check as an automated regression test. Make it fail on the old behaviour and pass on the fixed behaviour.

For a one-off copy change, reading the page may be enough. For duplicate payments or a broken retry, a durable test protects the next edit.

The next pull request should say both what the suite proved and what somebody did in the running app. That gives [the person reviewing it](/blog/how-to-review-an-ai-generated-pull-request) something concrete to repeat.

## is a screenshot enough for a UI change?

It can be enough to check a small copy edit. Use interaction evidence for controls, retries, focus or downloads. Put those actions in [the original brief](/blog/how-to-write-a-brief-for-a-coding-agent) so the next reviewer can repeat them.


[HTML version](https://wayari.com/blog/a-green-test-suite-never-opened-the-page)
