Agents / Skills

Grounding Check

Skill

Verifies that your AI feature's answers actually come from their sources: claims traced to the retrieved or provided material, inventions caught and classified, and the honest-refusal behaviour tested — the fluency-versus-faithfulness audit.

Best for

  • RAG and summarisation features that sound right more reliably than they are right
  • Quantifying hallucination before a stakeholder does it for you
  • Testing whether 'not in the sources' gets said when it should

What you give it

  • The feature's outputs with their source material (and the no-answer-exists cases for refusal testing)

What you get back

  • The faithfulness audit: claims decomposed and traced to sources — supported, unsupported, contradicted — with rates by claim type
  • The invention taxonomy: what gets made up (specifics like numbers and names, connective tissue, confident syntheses) and under which conditions
  • The refusal test results: whether the system says 'not found' when the sources genuinely lack the answer — or fluently invents one

How it works

  1. Decomposes outputs into atomic claims: each checkable statement separated, because sentence-level judgement lets inventions ride along with truths.
  2. Traces each claim to the sources: supported (with the span), unsupported (absent), contradicted (the source says otherwise) — specifics like numbers, names and dates checked exactly.
  3. Tests refusal deliberately: questions whose answers are absent from the sources probe whether honesty or fluency wins — the test most systems fail and most audits skip.
  4. Classifies the failures into the taxonomy that guides fixes: invented specifics, cross-source blending, over-synthesis, source-tone amplification.

Example

You: Audit our contract-assistant's answers before legal signs off on the rollout.

Result: The audit on 150 answer-source pairs: claims decomposed (1,100 total), 91% supported — but the 9% told the story legal needed: unsupported claims clustered in specifics (clause numbers cited that do not exist: 14 cases; dates shifted: 6) and in confident synthesis across documents (combining two contracts' terms into one answer: the worst class, 11 cases). The refusal test was the sharpest finding: on 30 questions whose answers are genuinely absent from the sources, the system answered anyway 19 times, fluently. Fixes applied (citation-constrained generation, the refusal rule with examples) and re-audited: supported rate 98%, refusals 28 of 30, the remaining failure modes documented for legal's conditions-of-use.

Limits — please read

  • Claim tracing at scale uses model assistance with human-validated sampling; the judge's own agreement rate is reported alongside.
  • Supported-by-sources is not the same as true — a faithful answer from a wrong source document is faithful; corpus accuracy is its own audit.
  • Rates reflect the tested distribution; the ugly-case oversampling (long sources, cross-document questions) is deliberate and stated.