Not a wall of green. What changed, what it costs, what needs a person.

Recurring runs are compared to the last one, so the report is a delta: regressed, fixed, still failing, added, dropped. Tests that flip between runs are quarantined instead of colouring the whole suite. And every job carries its cost, so spend is a column you can filter.

In the dashboard
Suite health · Reporting · Cost
Feature specs
54, 55, 56, 57, 71
Run delta · nightly smoke · 22 Sep vs 21 Sep

nightly smoke (seeded) · 41 tests · pass rate 87% (was 90%)

ChangeTest
regressedhotel-inventory-adjust
fixedrooming-list-export-xlsx
still failingparticipant-declined-card (3 runs)
quarantinedblock-request-partial (flipped 4 of last 6)

Blocked and cancelled runs are not verdicts and never count as flakiness.

The digest only speaks when something moved.

A trigger with a webhook attached posts this after each firing, and stays silent when the delta is empty. The Reporting view keeps the same comparison for any two runs of a trigger, with a pass-rate trend over the last N runs and a flakiness ranking by verdict flips.

Quarantine rules, exactly

  • Flip-flopping (pass and fail alternating over recent runs) puts a test in quarantine: it still runs and still reports, but it no longer turns the suite red.
  • Failing every run is not flaky, it is broken, and it still alarms.
  • Settling failing after quarantine re-alarms; settling passing releases it.
  • Blocked or cancelled runs are infrastructure outcomes, excluded from the flip count.
The Suite health view: every suite listed unhealthy first, regressed and fixed totals, cross-suite regressions, and the quarantine and muted-coverage lists.

One morning screen.

Suite health lists every suite unhealthy first, totals regressions and fixes, calls out tests that regressed in more than one suite, and shows the quarantine list and the tests PR planners are routing around. On-demand flaky analysis reads the recent runs and says whether the app, the agent or the environment is the likely cause, with a confidence and a suggested next step, and does nothing on its own.

Cost is on every job.

Provider, model, tokens and dollars are recorded per job, attributed correctly under concurrency. The Cost tab filters by kind, model, origin, lane and state, shows the trend, and is honest about gaps: an agent task that records no usage, a model without a price, a video render that costs nothing are shown as uncounted, not as zero.

The Reporting view: a trigger's runs compared, regressed and fixed tests listed, pass-rate trend.

Where it stops today

  • Deltas compare firings of the same trigger. Two ad-hoc runs of overlapping tests are not compared automatically.
  • Flaky analysis is an LLM reading run signals; it recommends, it does not quarantine or retire anything itself.
  • Cost is what the provider reports. Subscription-based harnesses (Claude Code on a person's plan) report tokens but no dollar price.
  • The health screen reflects suites that run on triggers; a project that only runs tests by hand has little to show here.