Exploratory testing, page by page, against a named list of checks.

A regression test proves one flow still works. Exploration asks a different question: for this page, what have we actually verified? The Explore screen answers it with a library of named checks grouped by category, run one page and one category at a time, with every result kept.

In the dashboard
Explore
Feature specs
38, 68
The Explore list view: every discovered page as a row with visit counts, last explored, and per-category pass and fail counts.

The check library.

Around four hundred checks in twenty-one categories, seeded as defaults and editable per project. Each check is a short instruction a browser agent can honestly carry out and a definition of pass, fail and skip. Checks a browser agent cannot perform (a load test, a server-side scan) ship disabled with the reason written on them; a project that has the means can turn them on.

One sweep job is one page against one category. Results accumulate per page, so the question "when did anyone last look at accessibility on the checkout page" has a date and a list.

  • Functional44
  • Accessibility44
  • Responsive27
  • Performance27
  • Visual & layout25
  • Security24
  • API & data20
  • Error handling19
  • i18n18
  • Resilience18
  • Auth17
  • State16
  • Content12
  • Usability13
  • Media11
  • Privacy11
  • Browser compat17
  • E2E regression14
Accessibility · / · one sweep job's results (this site, 23 Sep)
  • pass
    skip-navigation-links-work

    Skip to content moves focus to <main>; the link is the first tab stop.

  • fail
    text-contrast-meets-accessibility-requirements

    Secondary #7a7f88 text on white measured 4.02:1 at 14px, below 4.5:1. Affected breadcrumb labels and footer headings.

  • fail
    touch-target-sizes-meet-requirements

    Footer links render 20px tall; Home link 36×21.8px. Below the 24px minimum.

  • pass
    images-have-appropriate-alternative-text

    All 9 raster images carry descriptive alt text.

  • skip
    video-captions-are-available

    No video on this page.

Findings are written to be acted on.

These are real results from sweeping this marketing site before it shipped. Each failed check names the element, the measurement and the threshold, which is what let the two contrast and tap-target findings be fixed in one commit. The note is the agent's, not a template.

From a failed check you can ask for a regression test. An authoring run reproduces the finding in the browser and writes one focused test that fails on the current build and passes once the defect is fixed. It lands in the catalog disabled, in its own lane, so a sweep of a hundred pages never floods the main list.

Exploration from a pull request or an issue.

A watched pull request can carry an exploration of the pages its diff touched: the planner picks the pages and the categories, the results show on the PR's panel, and they do not decide the PR's verdict. An issue can generate an exploration brief, an implementation-agnostic description of what to probe plus an editable hints field, which becomes the agent's instructions.

Exploration runs also probe for defects on purpose: invalid and boundary input, cancel or refresh mid-flow, double submit, permission edges, judged against how the page ought to behave.

Where it stops today

  • A check is a judgement by an agent reading the page. Two runs of the same check can disagree; the history view is there so a person can see the pattern, not one verdict.
  • Some check methods are approximations. Zoom is simulated, keyboard traversal starts where the agent's last click left focus, and lazy-loading is inferred from network timing. Read a failed note before treating it as a defect.
  • A sweep boots an environment per project and runs about two jobs at a time on the shared worker pool, so a full 21-category sweep of a large app is an afternoon, not a minute.
  • SEO, PWA and analytics checks are off by default for internal dashboards; turn them on for a public site.