How a change flows from signal to shipped.
Seven stages, one queue, one set of workers. Every stage leaves evidence, and every stage where judgement matters waits for a person.
The seven stages.

-
A signal arrives
A pull request is opened or pushed, an issue is watched, a schedule fires, a signed webhook lands, or a person clicks Run. Every signal becomes jobs on one durable queue with leases and a visibility timeout.
-
The environment comes up
If the project declares one, a fresh copy of the app is built from the exact commit in a private namespace, patched, seeded with data and health-checked. Otherwise the run targets the configured URL.
-
The agent runs the tests
Plain-language tests, a real browser, prerequisites resolved from the catalog, screenshots and video captured as it goes, one forced retry before any failing verdict.
-
The verdict is posted
A label and one updating comment on the pull request, a retest comment on the issue, a Slack card, a webhook, a live event stream and the dashboard. Same payload everywhere.
-
Findings become work
A failed run becomes a filed bug with steps and video. An exploration finding becomes a regression test. A confirmed bug can be assigned to the code agent.
-
The fix is proven and handed over
The agent edits on an internal git host, boots the app from its branch, runs the proving test on a fresh environment, and a reviewer graduates the change into a real pull request, which is then watched like any other.
-
The documents follow
Merged pull requests update the requirements suite, contradictions become conflicts for a person, and the day's changelog is drafted behind a publish gate.
We test the platform with the platform.
Every pull request to QA Web Agent is watched by QA Web Agent. The dogfood project clones the PR, boots the whole platform (dashboard, API, workers, database, a mock git provider with a fictional repository) in a throwaway box, and runs the dashboard suite against it. Change detection, apply, test generation and release notes run on the same PR through the same queue. Explore sweeps run against the dashboard's own screens, including this site.
The pull request that added this site was tested that way: the planner authored three tests for it, the suite failed twice on real bugs (a redirect that lost its port, a canonical URL assertion that was wrong in the test), and the Explore sweeps produced the contrast, tap-target and reflow findings that were fixed before it merged.
services: mongo · api · test-worker · brd-worker · web (this dashboard)
mock-git (a fictional acme/demo-shop with issues, PRs, labels,
comments, a GitHub App, and an internal git host)
seed: SEED_PROJECTS=isnull
SEED_SLICES=releases,url-map,github-mock,trigger-history,
job-fixtures,issue-fixtures,pr-suite-noise,
code-challenges,brd-docs,vendor-jade,agentic-changes,
explore,site-discovery,brd-validation,output-video,
github-app,claude-agent,cost-history,brd-corrupt
auth: dev-token (the outer agent signs in with it)Three services, one queue.
An API, a browser worker and a document worker meet only at a durable queue. Foundational tests schedule first. Workers autoscale off queue depth and shrink to zero between bursts. A worker that silently wedges restarts itself; an infrastructure-blocked job is retried once as a fresh job; a job's scratch is removed the moment it ends and swept at boot. Every Azure resource is declared in terraform and every in-cluster object in versioned manifests, applied only after a validation step passes. The whole thing also runs on a laptop with Docker and no cloud account.
