Test the commit, not whatever happens to be deployed.
For any branch, pull request or commit the platform stands up a real copy of your application: cloned from your repository, built the way you say, patched for testability, seeded with data, health-checked, and driven over a private connection. When the run ends the environment is gone.
| Testing a shared staging URL | Testing an isolated environment |
|---|---|
| Whatever was deployed last, by anyone | Exactly the commit under review, pinned by SHA |
| One team's test data trampled by another's | Fresh seed per environment, or a project reset script between runs |
| Sign-in through the real identity provider | A patch swaps in a test-auth bypass; the real provider is untouched |
| Outbound email goes to real inboxes, or nowhere | A mail catcher runs beside the app and the agent reads it |
| A failed deploy blocks every tester | A failed boot blocks one job, with the app's own log attached |
You describe the app once.
The environment definition is part of the project: how to start the app, which port answers, what healthy means, how long to wait, which patches to apply, which uploads to place on the box, what memory, CPU and disk it needs. Three run modes cover most stacks: native install on a plain box, docker compose, or a prebuilt image.
Patches are ordered unified diffs applied to the throwaway clone, so you never fork your app for testing. When upstream moves and a patch no longer applies, the run fails fast with a machine-readable reason, and an agent task can repair the patch, prove a box boots with it, and resubmit the job.
enabled: true
mode: native (box image: sut-base)
ref: main # PR suites override with the head SHA
startup: npm ci && npx nx run-many -t build && npm run serve:all
port: 4200
healthcheck: GET /healthz → 200 within 900s
patches:
- jade-dev-auth.diff # B2C bypass for the test session
- jade-material-a11y.diff # expose Material control state to the a11y tree
uploads:
- jade-event.tgz → /opt/jade-event.tgz
- leonardo-data.json → /opt/leonardo-data.json
resources: memory 12Gi · cpu 2 · disk 15Gi
mailbox: enabled (Mailpit beside the app)
reset: scripts/reset-event.sh (exit 0 = clean)
What happens around the box.
- Private by default: each environment runs in its own namespace with default-deny networking and mutual TLS on the worker-to-app hop; only the assigned worker can reach it.
- Secure context for real auth: the worker reaches the app over a deterministic loopback port, so Web Crypto, MSAL and magic-link redirects behave as they do in production.
- Reuse when it is safe: a branch's environment stays warm for its next run after a liveness probe and, if the project defines one, a reset script that returns the data to seeded state.
- A ledger, always: no box exists without a record; a boot-time and periodic reconcile tears down anything the store no longer knows.
- Dependency cache: keyed by lockfile content, so the second boot of a monorepo skips the install.
- Watch it live: a gateway link from the Environments view lets a developer open a private box in their own browser while a run is on it.
Where it stops today
- The cluster backend runs on the platform's own Kubernetes estate. "Bring your own cluster" is not a supported story yet.
- A monorepo build is minutes, not seconds. Warm reuse and the dependency cache help; the first boot of a new commit does not get faster than the app's own build.
- Agent-triggered mid-run reset of a shared environment is switched off in production pending a durability fix; the between-runs reset script is what ships.
- Uploads and patches are copied into preview deployments from main at boot; a project created only on a preview starts empty.