Every verdict shows its work.

A pass or fail from an agent is only useful if someone who was not there can check it in a minute. So every run keeps the same set of evidence, and every link to it keeps working for months.

In the dashboard
Jobs · Video editing
Feature specs
07, 08, 30, 31, 66
A finished job's artifacts (API: GET /jobs/:id)
"state": "failed",
"summary": "Retry confirmed: /app returns ERR_CONNECTION_REFUSED …",
"artifact_urls": {
  "test.md":              "…/artifacts/0fbfd3fb…/test.md?exp=…&sig=…",
  "script.ts":            "…/artifacts/0fbfd3fb…/script.ts?exp=…&sig=…",
  "session.webm":         "…/artifacts/0fbfd3fb…/session.webm?exp=…",
  "session-1.webm":       "…/artifacts/0fbfd3fb…/session-1.webm?exp=…",
  "output.mp4":           "…/artifacts/0fbfd3fb…/output.mp4?exp=…",
  "tool-calls.jsonl":     "…/artifacts/0fbfd3fb…/tool-calls.jsonl?exp=…",
  "browser-console.log":  "…",
  "browser-network.log":  "…"
},
"screenshots": [
  { "title": "Step 6 — /app redirect refused", "page_url": "http://127.0.0.1:35075/app", "url": "…" },
  { "title": "Final state — FAIL", "auto": true, "url": "…" }
],
"usage": { "provider": "hermes", "model": "gpt-5.6-luna", "total_tokens": 57289, "cost_usd": 0.41 }

What every run leaves behind.

  • Screenshots the agent chose to take, each with a title in its own words, plus an automatic capture on every new URL and a final-state shot labelled with the verdict.
  • Whole-session video, per tab. Recorded from the browser context, remuxed so it seeks, and out of the agent's reach: there is no tool to stop or trim it.
  • One captioned walkthrough, the tabs merged into a single video with the agent's tool calls burned in as subtitles and long static stretches trimmed to a few seconds.
  • The transcript, system prompt, every turn, every tool call and result, provider and model, streamed live while the run is still going.
  • Browser console and network logs, untruncated, on any non-pass verdict, rendered as tables on the job.
  • Cost. Model, tokens and dollars on the job record.
The Video editing view: a rendered walkthrough with caption and static-trim toggles and the tool-call timeline.

Links that do not die.

Every artifact URL is signed by the platform and valid for months, and the API re-signs on every read. A Slack message from last quarter, a GitHub comment on a closed PR, a link pasted into a ticket: they still open. Screenshots posted to GitHub are embedded inline, not attached as expiring blobs.

Evidence goes where the reader is.

The same screenshots and recordings appear on the job in the dashboard, in the PR comment, on the filed issue, in the Slack card, in the code-change review, and in the changelog draft (only shots from passing checks are offered there).

Where it stops today

  • Video is per browser tab. A flow that opens a new window produces a second file (session-1.webm); the walkthrough render stitches them by frame change, which is a heuristic.
  • The static-trim in the walkthrough uses a freeze detector; a page that animates constantly will not be trimmed.
  • Console and network logs are captured on non-pass verdicts only, to keep passing runs light.
  • Screenshots are what the agent chose to capture plus the automatic ones; a step the agent did not think worth a picture has the video and the transcript, not a still.