Best for
- Validating a change before you call it done
- Separating a flaky test from a real failure
- Getting a plain pass/fail report instead of terminal noise
What you give it
- The change to check (it reads git diff)
- Your test commands (it reads your package scripts or a short TESTING.md)
- Anything that must not be run: for example database resets
What you get back
- The commands it ran and the pass/fail counts
- Each failure with the file and the likely cause
- A clear label: Implemented, Tested, Partially tested, or Blocked
How it works
- Looks at what changed (git diff --stat) to choose the relevant checks.
- Runs type-check, then unit and integration tests; browser tests only when asked.
- For every failure it reads the output, names the likely root cause, and re-runs once to tell flaky from real.
- Never edits files and never runs destructive commands.
- Reports honestly: a skipped or failing check is never described as passing.
Example
You: I changed the invoice export. Please validate it.
Result: Ran type-check (clean), unit tests 212/212, integration 48/50. Two failures in export.test.ts: timezone off by one in the date column (cause: local time used instead of UTC). Re-ran: same failure, so real, not flaky. Browser tests not run (you did not ask).
Limits — please read
- It is only as good as the tests you have.
- It reports; it does not fix.
- Long browser suites can be slow: ask for them when UI flows changed.