The pattern went from a viral post to a Google specification in ten weeks, and produced a pile of build-guides along the way. I went looking for a published comparison of which model to generate one with, what that costs, and how it fails, and did not find one — if you know of one I would genuinely like to read it. So this is the measurement I ran instead, including the part where the benchmark turns on itself.
Measured 2026-08-17. Model versions and prices move; these are snapshots, and the rate book is versioned in the repo alongside the scorecards. Everything here is a snapshot, not a standing claim.
One document in, one structured page out: YAML frontmatter plus a markdown body, produced
by a four-step loop (survey, extract, relate, compose). Fourteen candidate models generated the same
16 gold-labelled documents through the same loop and were scored by an
identical judge prompt (sha256 12d7ddaa… on every arm) against a
seven-criterion weighted rubric. The judge is anthropic:claude-sonnet-4-5-20250929 — a different
vendor from every arm it scores, so no model grades its own homework. Cost and latency come from a
per-call ledger, not from a price list.
Before any of the numbers below: that judge has never been calibrated against a
human. No kappa exists, it carries 0.5 of the rubric weight, and every
scorecard behind this page reports citable: false. What that does and does not invalidate is
in What this does not tell you, and you should read it before you quote anything
here.
gpt-5.5 scores 0.9073 at $0.2152 per document.
gpt-5.6-luna scores 0.8953 at $0.0153 —
14.1× cheaper, 1.4× faster, and not
separated by this measurement.
Paired across the same 16 documents the difference is
+0.0120 with a 95% confidence interval of
[-0.0353, +0.0593]. The interval spans zero. With one trial per document
and an uncalibrated judge, that is not a narrow victory — it is a difference this instrument cannot
resolve.
That is the practical finding: on this task, above a certain floor, the quality axis flattens and the cost axis does not. 6 of the fourteen arms have comparable means and land within 0.06 of the top score, and the spend across those 6 varies by 44×. If you picked your generator by reputation, you probably paid a multiple for a difference this benchmark cannot resolve.
Cost is a log scale — the spread is two orders of magnitude. Up and to the left is better. Models that failed to produce a usable page on every document are drawn hollow and faint: they are not viable at any price. 7 arms crashed on documents the others completed, so their cost is a proven floor, not a total — those carry a right-pointing whisker, and their true position lies somewhere to the right of the marker.
Sorted by rubric score. Two rules apply to every row with no exceptions: the document gate is the number of documents scoring ≥ 0.75 divided by 16 — the whole corpus, never the survivors — and cost per document is total ledger spend divided by 16. An arm that cannot clear the gate consistently is not viable at any price, which is why the cheap end of this table is mostly unusable.
| Model | Score | Doc gate | $/doc | Score per $ | Latency | Output tok | Docs scored | Empty |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5premium baseline | 0.9073 | 100% | $0.2152 | 4 | 71.7s | 117,261 | 16/16 | — |
| gpt-4.1extract gpt-4.1-mini | 0.9066 | 100% | $0.0641 | 14 | 60.8s | 98,369 | 16/16 | — |
| kimi-k2.611 of 16 generations returned an empty body · scored on 5 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · one OpenRouter route; rate includes a 39% promotional discount, non-promotional routes run ~39% higher on output | 0.9023 † | 25% | ≥$0.0554 | — | 206.2s | 56,556 | 5/16 | 11 EMPTY |
| gpt-5.4-nano | 0.9010 | 100% | $0.0154 | 59 | 52.9s | 126,731 | 16/16 | — |
| gpt-5.6-lunatied with the top arm at 1/14th the cost | 0.8953 | 94% | $0.0153 | 59 | 52.6s | 133,300 | 16/16 | — |
| qwen3-32b3 failed on malformed output or schema validation · scored on 13 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · the only arm given more than one attempt — two earlier routes failed on 2026-08-10, this is the third; slowest arm | 0.8745 † | 81% | ≥$0.0063 | — | 306.7s | 79,162 | 13/16 | — |
| qwen3.7-flashone OpenRouter route · rate triples above 32,000 input tokens | 0.8740 | 81% | $0.0049 | 178 | 91.8s | 282,766 | 16/16 | — |
| gpt-5-nano7 of 16 generations returned an empty body · scored on 9 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · scheduled shutdown 2026-12-11 | 0.8708 † | 56% | ≥$0.0084 | — | 91.0s | 128,413 | 9/16 | 7 EMPTY |
| gemini-3.1-flash-litefastest arm | 0.8683 | 94% | $0.0133 | 65 | 12.0s | 57,374 | 16/16 | — |
| seed-1.6-flash4 failed on malformed output or schema validation · scored on 12 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · one OpenRouter route · output rate more than doubles above 128,000 input tokens | 0.8505 † | 50% | ≥$0.0043 | — | 69.0s | 112,760 | 12/16 | — |
| gemini-2.5-flash-lite4 failed on malformed output or schema validation · scored on 12 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them | 0.8423 † | 62% | ≥$0.0061 | — | 34.8s | 143,871 | 12/16 | — |
| gpt-4.1-nano1 of 16 generations returned an empty body · 1 failed on malformed output or schema validation · scored on 14 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · retires 2026-10-23 | 0.7947 † | 62% | ≥$0.0037 | — | 22.9s | 33,885 | 14/16 | 1 EMPTY |
| gpt-4o-miniwidely deployed default | 0.7768 | 62% | $0.0055 | 141 | 31.6s | 43,127 | 16/16 | — |
| deepseek-v4-flashNO MEASUREMENT - every document failed at the provider (HTTP 404, no route available for this model id) · one OpenRouter route, DeepInfra fp4 — quantisation is a real variable here | 0.0000 † | 0% | ≥$0.0000 | — | 0.0s | 0 | 0/16 | — |
† This arm's score is a mean over a self-selected subset. It crashed on documents the other arms completed, so the surviving documents are an easier corpus than the one everybody else was measured on. These means are not comparable to the rest of the table in either direction — the crashes may have removed the hardest documents or the easiest, and nothing here distinguishes the two.
≥ The run crashed mid-document, so a total cost does not exist. The figure shown is what the ledger proves was already billed, divided by all 16 documents — a floor, never a total. Score per $ is — rather than a number, because a ratio built on a floor would read as a measurement.
Output tok counts only documents that completed. Crashed documents burned tokens
too — a further 3,049 for gemini-2.5-flash-lite and 1,174 for
seed-1.6-flash — and produced nothing.
Every one of these was invisible until the models actually ran. None of them is derivable from a rate card, a context-window number, or a public leaderboard.
gpt-4o-mini emitted 43,127 output tokens across the corpus and
gpt-4.1-nano emitted 33,885, against
gpt-5.6-luna's 133,300 on the identical documents — roughly a
third. They are cheaper partly because the artefact is thinner, and they miss the document
gate 38% and 38% of the time. Cost per token
is not cost per usable page.
gpt-5-nano lists at a fraction of the recommended model's rate and emitted
128,413 output tokens — 1.0× as
many as gpt-5.6-luna, all billed at the output rate. It ended up the
lowest-scoring OpenAI arm at 0.8708 and among the slowest at 91.0s per
document. A published price is an arithmetic input, not a forecast.
gemini-2.5-flash-lite advertises a million-token window and failed on two of sixteen
documents — for two different reasons. One was size: JSON truncated at character 211,025 on
the largest document. The other was not: on a roughly 1,000-token blog post it returned well-formed
JSON using a claim_type value outside our schema's enum, and our validator rejected it.
Call that the model's fault or our schema's — window size explains neither.
Both appear in the public price book. Both accept requests. Neither returned a structured artefact, which is the entire job. Both refused models were probed through one broker, in one mode, on 2026-08-07. Neither was tested against its vendor's own API, so this is a statement about those routes on that date and not about the models in general.
ling-3.0-flash — the cheapest row in the book'does not support feature: structured-outputs'. response_format is absent from supported_parameters on both its endpoints.
qwen3.5-flash — answers, but not with an objectIn JSON mode it returns the bare float -1.0000000000000002e+308 instead of an
object. Five attempts out of five, two system prompts, two user messages,
finish_reason: "stop", billed in full. A single endpoint, so there is no other route
to try.
The rule these produced, in order: a rate row proves arithmetic, not availability — then availability does not prove usability — then an endpoint that answers is not an endpoint that works.
This is the one that would have shipped, and it is the reason a benchmark has to check the artefact and not just the exit code.
19 generations returned a well-formed envelope with an EMPTY body (kimi-k2.6 11, gpt-5-nano 7, gpt-4.1-nano 1), across 11 distinct documents. They are not zero-byte files - the frontmatter parses, which is exactly why a downstream indexer would accept them. Between them they burned 417,388 output tokens (7,004 to 28,714 each) producing nothing. In the previous run these were scored as SUCCESSFUL documents; a guard now rejects an empty body before it can be scored, which is why they appear here as errors and why the affected arms report a lower completed-document count rather than a quietly lower score.
The previous version of this benchmark scored every one of them as a successful
document. The row said ok, the error count said zero, and the only trace was a
token count of zero on a field nothing was checking. A pipeline watching exit codes would have
indexed all 19 and billed for them.
The fix is a guard that rejects an empty body before it can be scored, and it belongs in the worker as well as the benchmark — a benchmark-only fix leaves the product broken. That is why these appear here as errors rather than as quietly low scores, and why the affected arms report fewer completed documents instead of a slightly worse mean. Read the gate column, not the mean: an arm that fails a document is not scored on it, so failing more can look like scoring higher.
The section that earns the rest of the page. Everything below is a limit we found in our own instrument, and the first item cost us a headline number.
faithfulness score does not measure faithfulness. It is the harmonic mean
of judged precision (are the summary's claims supported?) and judged recall (are the source's salient
facts present?). Precision sits on the ceiling — mean 0.9640, exactly 1.00 on
141/174 cells — so recall does all the moving. The criterion therefore
tracks recall at Spearman +0.99 and precision at only
+0.35, and 23 cells scored
below 0.50 while fabricating nothing at all. It carries 0.25 of the rubric weight,
so it shapes every mean on this page. A low score there means the summary left things out. It does not
mean the model made things up. We left the weights alone rather than silently redefine the unit every
historical number was measured in.must_not_contain
scan for 28 author-curated landmine phrases — plausible-but-false statements a
fabricating generator might emit. Not one was tripped anywhere. But 16 documents
with zero failures bounds the true rate only at ≤17.1%
(exact one-sided 95% Clopper–Pearson). That is not evidence of a general absence of fabrication.citable: false. The deterministic criteria, the cost ledger and the latency numbers are
unaffected.coverage is also a
completeness measure and saturates at 1.0 on the strong arms. So 0.50 of the weight sits on two
correlated completeness measures, and 0.10 on the only deterministic hallucination check.gpt-5.6-luna versus qwen3.7-flash differs by
+0.0214. Both intervals span zero. Decide on cost, speed or
operational risk, because quality is not deciding it for you.Every figure on this page is emitted by a build script from committed run scorecards, and a verifier recomputes each one from the raw judged cells before publication — including the statistics, which are derived from the per-document rows rather than copied from a stored aggregate.
Prose is the weak point and it is worth being explicit about it. A verifier that checks numbers cannot check a sentence, and an earlier draft of this page carried a whole narrative from a previous run underneath a correct table. The generator now rebuilds any sentence that restates a figure, and refuses to write a ledger where a superseded number survives in prose — but a claim whose premise has quietly stopped being true is still something only a reader can catch.
Supersession. This publication SUPERSEDES an earlier one measured on a different corpus. The earlier numbers do not carry over and should not be compared to these. What changed:
Said plainly because a page arguing that unverifiable numbers are not evidence does not get to edit its own history quietly. If you saw the earlier figures, they were measured on a corpus that included two documents this repository does not contain, which is why they are not reproducible here and are not repeated.
| Arms | 14 models, 16 documents each |
| Judge | anthropic:claude-sonnet-4-5-20250929 — cross-vendor, no arm judges itself |
| Judged criteria | faithfulness, coverage (0.5 of rubric weight) |
| Rubric weights | faithfulness 0.25, coverage 0.25, metadata_accuracy 0.2, anti_hallucination 0.1, format_compliance 0.1, citation_validity 0.05, token_efficiency 0.05 |
| Trials | 1 per document per arm |
| Metered spend | $6.6882 across all 14 arms (two are floors) |
| Judge calibration | never computed — kappa gate 0.6 unmet, citable: false |