80.3% raw. 94–98% shipped.
The difference is what our verification stack does — and this page is how we measured it, including the part that looks bad.
Published 2026-08-04. If a citation in your memo doesn't hold up, you find out in the room. Every vendor in this category claims their citations are reliable; as far as we can find, none publishes a measured rate for its own product. This is ours.
The short version
Across 1,166 cited sentences in six long-form reports, judged by a model from a different provider than the one that wrote them:
- 80.3% citation integrity from raw generation, with every repair mechanism switched off.
- 94.1%, 96.6%, 98.0% on the three reports produced by our full production chain — the reports you can open and check right now.
- The difference — roughly 14 to 18 points — is what our verification stack contributes. That gap is the product.
What was measured
Six research tasks drawn from a public deep-research benchmark, each run through discovery, source interrogation, and drafting at roughly 16,000 words over about 50 sources — but with the entire repair layer disabled: no verification gate during drafting, no post-draft audit loop, no repair pass. That isolates generation-time accuracy: what the model produces before any of our defenses touch it.
Every cited sentence was then read by gpt-5.5 alongside the sources it cites and classified supported, partially supported, or unsupported. The judge comes from a different provider than the drafting model on purpose: a model grading its own writing scores it too kindly. Full rubric on the methodology page.
Results, per report
The tasks: Japanese demographics and consumption to 2050; the investment philosophies of Buffett, Munger and Duan Yongping; how the wealthiest sovereign funds invest; machine learning in asset allocation; quantitative trading strategies; and auction theory. Five of the six are finance or economics — a narrow sample, and worth knowing before you generalize from it.
| Report | Integrity rate | Fully supported only | Cited sentences | Supported / partial / unsupported |
|---|---|---|---|---|
| Report 1 | 84.2% | 72.9% | 199 | 145 / 45 / 9 |
| Report 2 | 82.8% | 69.4% | 186 | 129 / 50 / 7 |
| Report 3 | 78.1% | 65.5% | 203 | 133 / 51 / 19 |
| Report 4 | 78.7% | 73.4% | 188 | 138 / 20 / 30 |
| Report 5 | 79.6% | 67.6% | 216 | 146 / 52 / 18 |
| Report 6 | 78.2% | 67.8% | 174 | 118 / 36 / 20 |
| Pooled | 80.3% | 69.4% 95% CI 66.7–72.0% |
1,166 | 809 / 254 / 103 |
Two columns, because the convention matters: our published rate gives partial credit (supported counts 1, partially supported 0.5), while the stricter column counts only fully-supported sentences. On this sample that is an 11-point difference. Counting sentences that are not unsupported, the figure is 91.2%. All three are the same data; only a stated convention is checkable.
The comparison that matters
Those six reports are not what we sell. They are the raw material. Our production chain adds generation-time grounding, a verification gate that blocks unsupported specifics as the draft is written, and a bounded repair pass. The three sample reports on this site went through it:
| Report | Measured integrity | Length |
|---|---|---|
| AI data-center electricity demand through 2030 | 94.1% (95% CI 89–97%) | 13,384 words |
| The decipherment of Linear B | 96.6% (95% CI 92–99%) | 8,872 words |
| History of the B-52 bomber | 98.0% (95% CI 90–100%) | 3,985 words |
Publishing only the second table would be true and useless — an unfalsifiable number with nothing to compare it against. The pair is the honest form: it shows what the machinery is worth, and it means a reader can judge the claim rather than take it.
What this measurement does not show
- It is self-run. We built the instrument, chose the sample, and paid for the judge. The judge is independent of our drafting model, but the study is not independent of us. We would rather someone else ran it — see below.
- It is not a competitor comparison. No other product was tested. We deliberately did not benchmark rivals on our own harness; self-run competitive benchmarks deserve the skepticism they get.
- The sample is six reports, and it leans finance. 1,166 cited sentences is enough for a tight interval on these reports; it is not enough to characterize every topic, language, or domain, and five of the six tasks come from finance and economics.
- The judge is a model. Classifying whether a paraphrase is supported involves judgment and carries its own error rate. The deterministic parts of our proof record — verbatim spans and numeric grounding — do not.
- It measures faithfulness to sources, not the truth of sources. A sentence that accurately reports a wrong source counts as supported here. Source quality is a separate problem, handled separately.
- The confidence interval is an approximation for a partial-credit rate, which is not a clean pass/fail outcome. Treat it as calibration.
Check it, or repeat it
Every report we publish ships a per-sentence proof record, and a small script re-checks it on your machine against the live cited sources — no account, no dependence on us.
Want to run this yourself?
We will hand over the protocol, the task set, and the raw audit files — no conditions. If you evaluate research tools professionally, that is the offer: info@siderealintelligence.com. An independent measurement of this category does not exist yet; we would rather help someone write it than keep being the only ones grading our own work.
We re-run this measurement when the pipeline changes in ways that could move it, and date every revision. The figures above are from 2026-08-04.