How we measure citation integrity
Every report we deliver displays a measured citation-integrity figure with a confidence interval. This page defines that number: how it is produced, what it proves, and — just as important — what it does not.
The unit of measurement: the sentence
A report's integrity rate is computed per sentence, not per document. Every cited sentence in the finished report is read by an auditing model alongside the source material it cites, and classified: supported (the cited evidence substantiates what the sentence says), partially supported (the citation is relevant but the sentence overreaches it), or unsupported (the cited evidence does not substantiate the sentence). The integrity rate is a half-credit-weighted proportion: supported sentences count 1, partially supported count 0.5, unsupported count 0. We state the convention because it materially changes the number — on a recent six-report sample the same audits read 80.3% half-credit and 69.4% counting only fully-supported sentences. Both are honest; only a stated convention is checkable. Every per-report figure we display uses the half-credit rate, and the audit sidecar downloadable with each report carries the raw supported / partial / unsupported counts so you can recompute it any way you prefer.
The judge is independent of the writer
The model that audits a report comes from a different provider than the model that drafted it. A drafting model grading its own work exhibits self-preference bias — it recognizes its own phrasing and grades it leniently — so we route the audit to an independent model family and accept the stricter verdicts that result. When a report's rate sits at a decision boundary (for example, the refund threshold below), flagged sentences are additionally re-judged by a second cross-family adjudicator, and both the strict and adjudicated rates are recorded.
The confidence interval
A rate computed over a finite number of sentences carries sampling uncertainty, so every figure we display ships with a 95% Wilson confidence interval. (The Wilson interval is exact for pass/fail outcomes; because our rate is half-credit-weighted, treat the band as calibration rather than a strict inferential claim.) A report showing "94.1% (95% CI 88–97%)" is telling you both the measurement and how much to trust it. We display the interval even when it is unflattering.
The proof sidecar: checkable without trusting us
Alongside the judged rate, every report ships a deterministic per-sentence proof record. No model is involved in producing or checking it — it is a pure function of the report and its evidence. Each cited sentence is classified as:
- verbatim — the sentence anchors character-for-character to a span in a cited source;
- grounded — every checkable specific in the sentence (numbers, years, quantities, named entities) is present in its cited sources' evidence;
- attributed — the sentence is paraphrase carrying no checkable specifics, resting on the citation and the judged audit;
- unverified — a specific could not be located in the cited evidence. These are counted and shown, not hidden.
On any public report page, the proof rail shows this record sentence by sentence, with the supporting quotes. The raw sidecar is downloadable as JSON from every public report, and a small standalone script re-checks it against the published page and the live cited sources — it runs on your machine, with no account and no dependence on our goodwill.
What the number does not prove
Stated plainly, because the failure mode of this category is overclaiming:
- It is a measurement, not a guarantee. A 95% rate means roughly one cited sentence in twenty did not fully clear the strictest reading of its evidence. We show you which ones.
- It measures faithfulness to sources, not the truth of sources. If a cited source is wrong, a sentence faithfully reporting it still counts as supported. Source quality is addressed separately — by interrogating each source's origin, audience, and bias, and by flagging corroborations that trace to a single shared origin.
- Judged classifications of paraphrase involve model judgment and carry a residual error rate of their own; the deterministic proof classes (verbatim, grounded) do not.
- These figures are self-measured: we built the instrument and we run it, using an independent judge model and the methodology on this page. We have not yet published an externally benchmarked category study; when we publish one, it will appear here first with its full protocol, sample, and residual error rate. Until then, the only numbers we show are per-report measurements.
The floor policy
Integrity governs delivery, not just display. A commissioned report measuring below our target earns one bounded repair pass and a re-audit. If it still measures below target, it ships with an explicit disclosure banner. If it measures below the floor, the run fails and you are not charged — failure is free. The decision trace for every such call is recorded with the run.
The fastest way to evaluate this methodology is against a real document: open a sample report, read its proof rail, and run the verifier yourself.