Issues that get re-tested until the evidence says fixed.
Two directions. Watch an issue someone reported, and the agent reproduces it, writes a repro test and retests on every commit that mentions the number. Or start from a failed run, and one click files a GitHub issue with the steps, the screenshots and the video, already watched, so a silent fix later shows up as a pass.
- Issue #101 watchedAn explore run reads the report, reproduces it in a browser, and authors a repro test that fails on the current build. The issue gets QA-WA-Validated.
- Commit references #101The sweep sees a new commit mentioning the issue and re-runs the repro test against an environment at that commit. Still failing; the comment says so.
- Another commit, now passingThe repro test passes. One more confirmation retest is scheduled before anyone is told it is fixed.
- ConfirmedThe confirmation passes too. The issue is marked verified on the platform. Closing it stays with a person.
What a watched issue's life looks like.
The repro test is the anchor. It is an ordinary catalog test that fails while the bug exists, so it does double duty: proof today, regression coverage forever. It lands disabled for review and is enabled when a person agrees it captures the report.
Every retest posts one comment, and a project can choose an aggregate comment that updates in place instead.
failed Still failing: quantity field accepts -1
Automated issue tracking (retest run) by qa-web-agent.
Triggered by commit: 68a065c
Step 3 · Entered -1 in the cart quantity field and pressed Update. The line total shows −$120.00 and the order total went negative. Expected a validation message and an unchanged total.
2 screenshots · session video
Cart accepts a negative quantity and produces a negative order total
QA-WA-Bug QA-WA-NeedsReview · type: Bug
Summary. On /cart the quantity input accepts -1; the line total and order total go negative and Checkout stays enabled.
Steps to reproduce.
- Open /products/widget-alpha and add one to the cart.
- On /cart set quantity to -1 and press Update.
- Observe the totals.
Evidence. final-state screenshot · session video · covering test: seeded-demo-shop-discount-check
AI-generated content may be incorrect.
Labels record who found it and what the agent thinks.
| Label | Set when |
|---|---|
| QA-WA-Bug | the agent filed the issue from a failed run |
| QA-WA-NeedsReview | a person has not yet confirmed the report |
| QA-WA-Validated | the agent reproduced the reported behaviour |
| QA-WA-Invalidated | the agent could not reproduce it |
| QA-WA-Non-UI | the report is not browser-testable; no run is burned |
The Tracking view rolls this up: coverage rings of open versus watched issues and bugs, how many need a person, and filings per day.

Where it stops today
- The agent never closes an issue and never reopens one. It posts still failing or now passing and marks the platform's record verified; the repository stays yours.
- A repro depends on the report. A vague issue produces a vague repro test; the explore brief it generates is editable before the run so a person can sharpen it.
- Retests key off commits that reference the issue number. A fix that never mentions #101 is only caught when the repro test runs in a scheduled suite.
- Filing needs a token or GitHub App with Issues write access. Without it the platform records posting skipped rather than failing the run.
Questions people ask
Does the agent close issues?
No. It reproduces, retests on every referencing commit and posts still failing or now passing. Closing stays a human decision.