Measurement

80.3% raw. 94–98% shipped.

The difference is what our verification stack does — and this page is how we measured it, including the part that looks bad.

Published 2026-08-04. If a citation in your memo doesn't hold up, you find out in the room. Every vendor in this category claims their citations are reliable; as far as we can find, none publishes a measured rate for its own product. This is ours.

The short version

Across 1,166 cited sentences in six long-form reports, judged by a model from a different provider than the one that wrote them:

  • 80.3% citation integrity from raw generation, with every repair mechanism switched off.
  • 94.1%, 96.6%, 98.0% on the three reports produced by our full production chain — the reports you can open and check right now.
  • The difference — roughly 14 to 18 points — is what our verification stack contributes. That gap is the product.

What was measured

Six research tasks drawn from a public deep-research benchmark, each run through discovery, source interrogation, and drafting at roughly 16,000 words over about 50 sources — but with the entire repair layer disabled: no verification gate during drafting, no post-draft audit loop, no repair pass. That isolates generation-time accuracy: what the model produces before any of our defenses touch it.

Every cited sentence was then read by gpt-5.5 alongside the sources it cites and classified supported, partially supported, or unsupported. The judge comes from a different provider than the drafting model on purpose: a model grading its own writing scores it too kindly. Full rubric on the methodology page.

Results, per report

The tasks: Japanese demographics and consumption to 2050; the investment philosophies of Buffett, Munger and Duan Yongping; how the wealthiest sovereign funds invest; machine learning in asset allocation; quantitative trading strategies; and auction theory. Five of the six are finance or economics — a narrow sample, and worth knowing before you generalize from it.

Report Integrity rate Fully supported only Cited sentences Supported / partial / unsupported
Report 1 84.2% 72.9% 199 145 / 45 / 9
Report 2 82.8% 69.4% 186 129 / 50 / 7
Report 3 78.1% 65.5% 203 133 / 51 / 19
Report 4 78.7% 73.4% 188 138 / 20 / 30
Report 5 79.6% 67.6% 216 146 / 52 / 18
Report 6 78.2% 67.8% 174 118 / 36 / 20
Pooled 80.3% 69.4%
95% CI 66.7–72.0%
1,166 809 / 254 / 103

Two columns, because the convention matters: our published rate gives partial credit (supported counts 1, partially supported 0.5), while the stricter column counts only fully-supported sentences. On this sample that is an 11-point difference. Counting sentences that are not unsupported, the figure is 91.2%. All three are the same data; only a stated convention is checkable.

The comparison that matters

Those six reports are not what we sell. They are the raw material. Our production chain adds generation-time grounding, a verification gate that blocks unsupported specifics as the draft is written, and a bounded repair pass. The three sample reports on this site went through it:

Report Measured integrity Length
AI data-center electricity demand through 203094.1% (95% CI 89–97%)13,384 words
The decipherment of Linear B96.6% (95% CI 92–99%)8,872 words
History of the B-52 bomber98.0% (95% CI 90–100%)3,985 words

Publishing only the second table would be true and useless — an unfalsifiable number with nothing to compare it against. The pair is the honest form: it shows what the machinery is worth, and it means a reader can judge the claim rather than take it.

What this measurement does not show

Check it, or repeat it

Every report we publish ships a per-sentence proof record, and a small script re-checks it on your machine against the live cited sources — no account, no dependence on us.

Want to run this yourself?

We will hand over the protocol, the task set, and the raw audit files — no conditions. If you evaluate research tools professionally, that is the offer: info@siderealintelligence.com. An independent measurement of this category does not exist yet; we would rather help someone write it than keep being the only ones grading our own work.

We re-run this measurement when the pipeline changes in ways that could move it, and date every revision. The figures above are from 2026-08-04.