Best for
- RAG and summarisation features that sound right more reliably than they are right
- Quantifying hallucination before a stakeholder does it for you
- Testing whether 'not in the sources' gets said when it should
What you give it
- The feature's outputs with their source material (and the no-answer-exists cases for refusal testing)
What you get back
- The faithfulness audit: claims decomposed and traced to sources — supported, unsupported, contradicted — with rates by claim type
- The invention taxonomy: what gets made up (specifics like numbers and names, connective tissue, confident syntheses) and under which conditions
- The refusal test results: whether the system says 'not found' when the sources genuinely lack the answer — or fluently invents one
How it works
- Decomposes outputs into atomic claims: each checkable statement separated, because sentence-level judgement lets inventions ride along with truths.
- Traces each claim to the sources: supported (with the span), unsupported (absent), contradicted (the source says otherwise) — specifics like numbers, names and dates checked exactly.
- Tests refusal deliberately: questions whose answers are absent from the sources probe whether honesty or fluency wins — the test most systems fail and most audits skip.
- Classifies the failures into the taxonomy that guides fixes: invented specifics, cross-source blending, over-synthesis, source-tone amplification.
Example
You: Audit our contract-assistant's answers before legal signs off on the rollout.
Result: The audit on 150 answer-source pairs: claims decomposed (1,100 total), 91% supported — but the 9% told the story legal needed: unsupported claims clustered in specifics (clause numbers cited that do not exist: 14 cases; dates shifted: 6) and in confident synthesis across documents (combining two contracts' terms into one answer: the worst class, 11 cases). The refusal test was the sharpest finding: on 30 questions whose answers are genuinely absent from the sources, the system answered anyway 19 times, fluently. Fixes applied (citation-constrained generation, the refusal rule with examples) and re-audited: supported rate 98%, refusals 28 of 30, the remaining failure modes documented for legal's conditions-of-use.
Limits — please read
- Claim tracing at scale uses model assistance with human-validated sampling; the judge's own agreement rate is reported alongside.
- Supported-by-sources is not the same as true — a faithful answer from a wrong source document is faithful; corpus accuracy is its own audit.
- Rates reflect the tested distribution; the ugly-case oversampling (long sources, cross-document questions) is deliberate and stated.