A regression test is a markdown file.
You write what a person would do and what they should see. An agent does it in a real browser, works out the prerequisites on its own, and returns pass, fail or blocked with the evidence to prove it. There is no selector to maintain and no script to re-record when the UI moves.
This is a real test from the Jade catalog.
Level 9, tagged housing, block-request, approve. Notice what it does not contain: no CSS selectors, no waits, no login steps. The preconditions name the tag of the test that handles login, and the agent runs that test itself when it meets the sign-in wall.
Levels order the catalog: foundational tests (login, data setup) sit at low levels and schedule first; a suite run walks up from there. Tags carry capabilities too: a tag can grant the agent authoring rights, a mobile viewport, or access to your app's MCP servers.
# Group block request: full allocation → Approved
Confirms an organizer can allocate the full requested rooms to a
pending block request and that its status becomes Approved.
## Preconditions
- Authenticated as an organizer. If redirected to a sign-in form,
use list_tests({ tags: ["login"], max_level: 1 }) and run_test
the first result.
- A pending room block request exists on the seeded event. If none,
use list_tests({ tags: ["block-request", "submit"] }) and run_test
the first result to create one.
## Steps
1. Open the seeded baseline event → Housing → Block Requests.
2. Select the Pending filter, click the pending request, note the
total requested rooms and the group name.
3. Click Add allocation until the room-nights are fully covered.
4. Click Publish changes.
## Expected result
The request shows status Approved and filled equals requested.What one run looks like on the clock.
Times from a representative Jade run. The retry at 01:58 is the rule that makes agent verdicts usable: a first FAIL or BLOCKED is never accepted on its own. A PASS is accepted at once.

- Job leased by a workerThe test's markdown is interpolated ({{ENV.BASE_URL}} and friends) and a fresh, isolated browser context opens.
- Login wallThe agent is redirected to sign-in. It calls list_tests with the login tag, finds the organizer login test, and runs it in the same session.
- Steps 1 to 4Navigates to Housing → Block Requests, opens the pending request, adds two allocations, publishes.
- First verdict: FAILThe status pill still reads Partial. The worker refuses the first FAIL, shows the agent the final screenshot and asks for one more genuine attempt.
- Retry: PASSA page refresh shows Approved. The agent reports PASS with the screenshot that proves it.
- Evidence uploadedScreenshots, both tabs' recordings, the transcript and a generated Playwright script are signed and attached to the job.
What the catalog gives you that a folder of scripts does not.
Prerequisites resolve at run time
Tests do not declare dependencies. When the agent meets a wall (sign-in, an onboarding wizard, an empty required dropdown) it searches the catalog for a test that clears it and runs that first, in the same browser. The transcript records which one it picked.
One catalog, three authors
Hand-written tests, tests generated from a requirements change, and tests the agent authored from an exploration finding all land in the same list. Agent-authored tests arrive disabled so a person reviews them before a suite can depend on them.
Invalidated, not deleted
A test that turns out to be wrong is marked invalid with a reason. Automated suites skip it, the history stays, and nobody wonders later why it vanished.
A generated script, for what it is worth
Every run also emits a Playwright script of what the agent did. A later run of the same unchanged test replays the script first and only wakes the agent if the replay breaks. Authoring runs never replay.

Where it stops today
- Agent runs are judgements, not deterministic scripts. Expect some noise; the retry rule, quarantine and per-run evidence exist because of it.
- Mobile is a viewport, not a device: a tag switches the window size, screenshots and video, but there is no touch emulation or device farm.
- A freshly authored test is not yet auto-validated in a clean environment before a person sees it; review is still load-bearing.
- The replayed Playwright script is an optimisation, not the source of truth. When it breaks the agent takes over; it is not meant to be exported and maintained elsewhere.
Questions people ask
How does the agent handle login screens and setup data?
It searches the test catalog for the prerequisite, for example a login or onboarding test, and runs it in the same browser session before continuing. No dependency graph to maintain.
Are the results deterministic?
No. An agent run is a judgement, and the same test can pass and fail across runs when the app or the agent behaves differently. The platform is built around that: one forced retry, quarantine for flip-flopping tests, and evidence on every run so a person can decide quickly.