Not a wall of green. What changed, what it costs, what needs a person.
Recurring runs are compared to the last one, so the report is a delta: regressed, fixed, still failing, added, dropped. Tests that flip between runs are quarantined instead of colouring the whole suite. And every job carries its cost, so spend is a column you can filter.
nightly smoke (seeded) · 41 tests · pass rate 87% (was 90%)
| Change | Test |
|---|---|
| regressed | hotel-inventory-adjust |
| fixed | rooming-list-export-xlsx |
| still failing | participant-declined-card (3 runs) |
| quarantined | block-request-partial (flipped 4 of last 6) |
Blocked and cancelled runs are not verdicts and never count as flakiness.
The digest only speaks when something moved.
A trigger with a webhook attached posts this after each firing, and stays silent when the delta is empty. The Reporting view keeps the same comparison for any two runs of a trigger, with a pass-rate trend over the last N runs and a flakiness ranking by verdict flips.
Quarantine rules, exactly
- Flip-flopping (pass and fail alternating over recent runs) puts a test in quarantine: it still runs and still reports, but it no longer turns the suite red.
- Failing every run is not flaky, it is broken, and it still alarms.
- Settling failing after quarantine re-alarms; settling passing releases it.
- Blocked or cancelled runs are infrastructure outcomes, excluded from the flip count.

One morning screen.
Suite health lists every suite unhealthy first, totals regressions and fixes, calls out tests that regressed in more than one suite, and shows the quarantine list and the tests PR planners are routing around. On-demand flaky analysis reads the recent runs and says whether the app, the agent or the environment is the likely cause, with a confidence and a suggested next step, and does nothing on its own.
Cost is on every job.
Provider, model, tokens and dollars are recorded per job, attributed correctly under concurrency. The Cost tab filters by kind, model, origin, lane and state, shows the trend, and is honest about gaps: an agent task that records no usage, a model without a price, a video render that costs nothing are shown as uncounted, not as zero.

Where it stops today
- Deltas compare firings of the same trigger. Two ad-hoc runs of overlapping tests are not compared automatically.
- Flaky analysis is an LLM reading run signals; it recommends, it does not quarantine or retire anything itself.
- Cost is what the provider reports. Subscription-based harnesses (Claude Code on a person's plan) report tokens but no dollar price.
- The health screen reflects suites that run on triggers; a project that only runs tests by hand has little to show here.