The pattern went from a viral post to a Google specification in ten weeks, and produced a pile of build-guides along the way. I went looking for a published comparison of which model to generate one with, what that costs, and how it fails, and did not find one — if you know of one I would genuinely like to read it. So this is the measurement I ran instead, including the part where the benchmark turns on itself.
Measured 2026-08-07, one arm re-measured 2026-08-10. Prices from a rate book verified 2026-08-06 to 2026-08-07. Model versions and prices move — everything here is a snapshot, not a standing claim.
One document in, one structured page out: YAML frontmatter plus a markdown body, produced
by a four-step loop (survey, extract, relate, compose). Fourteen candidate models generated the same
16 gold-labelled documents through the same loop and were scored by an
identical judge prompt (sha256 12d7ddaa… on every arm) against a
seven-criterion weighted rubric. The judge is anthropic:claude-sonnet-4-5-20250929 — a different
vendor from every arm it scores, so no model grades its own homework. Cost and latency come from a
per-call ledger, not from a price list.
Before any of the numbers below: that judge has never been calibrated against a
human. No kappa exists, it carries 0.5 of the rubric weight, and every
scorecard behind this page reports citable: false. What that does and does not invalidate is
in What this does not tell you, and you should read it before you quote anything
here.
gpt-5.5 scores 0.9366 at $0.2100 per document.
gpt-5.6-luna scores 0.9176 at $0.0147 —
14.3× cheaper, twice as fast, and not separated by this measurement.
Paired across the same 16 documents the difference is
+0.0190 with a 95% confidence interval of
[-0.0100, +0.0480]. The interval spans zero. With one trial per document
and an uncalibrated judge, that is not a narrow victory — it is a difference this instrument cannot
resolve.
That is the practical finding: on this task, above a certain floor, the quality axis flattens and the cost axis does not. 8 of the fourteen arms have comparable means and land within 0.06 of the top score, and the spend across those 8 varies by 48×. If you picked your generator by reputation, you probably paid a multiple for a difference this benchmark cannot resolve.
Cost is a log scale — the spread is two orders of magnitude. Up and to the left is better. Models that failed to produce a usable page on every document are drawn hollow and faint: they are not viable at any price. Two arms crashed on documents the others completed, so their cost is a proven floor, not a total — those carry a right-pointing whisker, and their true position lies somewhere to the right of the marker.
Sorted by rubric score. Two rules apply to every row with no exceptions: the document gate is the number of documents scoring ≥ 0.75 divided by 16 — the whole corpus, never the survivors — and cost per document is total ledger spend divided by 16. An arm that cannot clear the gate consistently is not viable at any price, which is why the cheap end of this table is mostly unusable.
| Model | Score | Doc gate | $/doc | Score per $ | Latency | Output tok | Docs scored | Empty |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5premium baseline | 0.9366 | 100% | $0.2100 | 4 | 84.8s | 113,055 | 16/16 | — |
| gpt-5.6-lunatied with the top arm at 1/14th the cost | 0.9176 | 100% | $0.0147 | 62 | 39.6s | 127,370 | 16/16 | — |
| qwen3.7-flashties luna, 3.3x cheaper, 2.6x slower · one OpenRouter route · rate triples above 32,000 input tokens | 0.9173 | 100% | $0.0044 | 208 | 102.1s | 248,993 | 16/16 | — |
| deepseek-v4-flashone OpenRouter route, DeepInfra fp4 — quantisation is a real variable here | 0.9102 | 100% | $0.0048 | 190 | 218.1s | 204,947 | 16/16 | — |
| gpt-4.1extract gpt-4.1-mini | 0.8979 | 100% | $0.0640 | 14 | 68.4s | 99,030 | 16/16 | — |
| qwen3-32bthe only arm given more than one attempt — two earlier routes failed on 2026-08-10, this is the third; slowest arm | 0.8896 | 94% | $0.0082 | 108 | 386.5s | 116,028 | 16/16 | — |
| gemini-3.1-flash-litefastest arm | 0.8816 | 100% | $0.0131 | 67 | 11.5s | 55,950 | 16/16 | — |
| gemini-2.5-flash-lite2 crashes | 0.8798 † | 81% | ≥$0.0068 | — | 32.3s | 155,863 | 14/16 | — |
| gpt-5.4-nano | 0.8767 | 88% | $0.0154 | 57 | 62.2s | 127,299 | 16/16 | — |
| seed-1.6-flashJSON top-level was a string; 2 crashes · one OpenRouter route · output rate more than doubles above 128,000 input tokens | 0.8098 † | 44% | ≥$0.0041 | — | 54.2s | 103,651 | 14/16 | — |
| gpt-4.1-nanoretires 2026-10-23 | 0.7993 | 75% | $0.0036 | 222 | 24.2s | 37,454 | 16/16 | — |
| gpt-4o-miniwidely deployed default | 0.7692 | 69% | $0.0053 | 145 | 26.5s | 38,827 | 16/16 | — |
| gpt-5-nano5 of 16 artefacts came back EMPTY · scheduled shutdown 2026-12-11 | 0.6740 | 62% | $0.0087 | 77 | 120.3s | 294,753 | 16/16 | 5 EMPTY |
| kimi-k2.611 of 16 artefacts came back EMPTY · one OpenRouter route; rate includes a 39% promotional discount, non-promotional routes run ~39% higher on output | 0.5259 | 31% | $0.0575 | 9 | 298.4s | 325,667 | 16/16 | 11 EMPTY |
† This arm's score is a mean over a self-selected subset. It crashed on documents the other arms completed, so the surviving documents are an easier corpus than the one everybody else was measured on. These means are not comparable to the rest of the table in either direction — the crashes may have removed the hardest documents or the easiest, and nothing here distinguishes the two.
≥ The run crashed mid-document, so a total cost does not exist. The figure shown is what the ledger proves was already billed, divided by all 16 documents — a floor, never a total. Score per $ is — rather than a number, because a ratio built on a floor would read as a measurement.
Output tok counts only documents that completed. Crashed documents burned tokens
too — a further 3,049 for gemini-2.5-flash-lite and 1,174 for
seed-1.6-flash — and produced nothing.
Every one of these was invisible until the models actually ran. None of them is derivable from a rate card, a context-window number, or a public leaderboard.
gpt-4o-mini emitted 38,827 output tokens across the corpus and
gpt-4.1-nano emitted 37,454, against
gpt-5.6-luna's 127,370 on the identical documents — roughly a
third. They are cheaper partly because the artefact is thinner, and they miss the document
gate 31% and 25% of the time. Cost per token
is not cost per usable page.
gpt-5-nano lists at a fraction of the recommended model's rate and emitted
294,753 output tokens — 2.3× as
many as gpt-5.6-luna, all billed at the output rate. It ended up the
lowest-scoring OpenAI arm at 0.6740 and among the slowest at 120.3s per
document. A published price is an arithmetic input, not a forecast.
gemini-2.5-flash-lite advertises a million-token window and failed on two of sixteen
documents — for two different reasons. One was size: JSON truncated at character 211,025 on
the largest document. The other was not: on a roughly 1,000-token blog post it returned well-formed
JSON using a claim_type value outside our schema's enum, and our validator rejected it.
Call that the model's fault or our schema's — window size explains neither.
Both appear in the public price book. Both accept requests. Neither returned a structured artefact, which is the entire job. Both refused models were probed through one broker, in one mode, on 2026-08-07. Neither was tested against its vendor's own API, so this is a statement about those routes on that date and not about the models in general.
ling-3.0-flash — the cheapest row in the book'does not support feature: structured-outputs'. response_format is absent from supported_parameters on both its endpoints.
qwen3.5-flash — answers, but not with an objectIn JSON mode it returns the bare float -1.0000000000000002e+308 instead of an
object. Five attempts out of five, two system prompts, two user messages,
finish_reason: "stop", billed in full. A single endpoint, so there is no other route
to try.
The rule these produced, in order: a rate row proves arithmetic, not availability — then availability does not prove usability — then an endpoint that answers is not an endpoint that works.
This is the one that would have shipped, and it is the reason a benchmark has to check the artefact and not just the exit code.
ok.kimi-k2.6 produced 11 of 16.
gpt-5-nano — from a major lab, not an exotic endpoint — produced
5 of 16. These are not zero-byte files, which is what makes them
dangerous: they carry valid YAML frontmatter and score 1.0 on schema validity, 1.0 on citation
validity and up to 0.89 on metadata accuracy, with
init_md_token_count: 0 and nothing whatsoever in the body. A downstream indexer accepts
them precisely because the envelope is well-formed.
Each burned 19,592 to 30,364 output tokens of output. So the failure is not merely silent, it is expensive and silent: real money spent, an empty-bodied artefact indexed, and no error, no warning and no non-zero exit to say so. The only trace is a token count of zero on a field nothing was checking.
One layer did catch it, and it is worth being precise about which. The row status lied and
the error count lied, but the aggregate document gate held —
suite_passed: false on both arms — and only because this benchmark happens to score the
artefact itself. A pipeline checking exit codes and token counts would have shipped all
16. The fix belongs in every generator of this kind: assert the artefact is non-empty
before recording the step as successful, in the worker and in the benchmark, because a
benchmark-only fix leaves the product broken.
Both arms ran one configuration, one trial per document, on 2026-08-07: gpt-5-nano on its listed deployment, kimi-k2.6 through a single OpenRouter route. This benchmark cannot separate model behaviour from route or broker behaviour, and neither arm was re-tested on an alternate endpoint.
The section that earns the rest of the page. Everything below is a limit we found in our own instrument, and the first item cost us a headline number.
faithfulness score does not measure faithfulness. It is the harmonic mean
of judged precision (are the summary's claims supported?) and judged recall (are the source's salient
facts present?). Precision sits on the ceiling — mean 0.9724, exactly 1.00 on
159/205 cells — so recall does all the moving. The criterion therefore
tracks recall at Spearman +0.99 and precision at only
+0.30, and 17 cells scored
below 0.50 while fabricating nothing at all. It carries 0.25 of the rubric weight,
so it shapes every mean on this page. A low score there means the summary left things out. It does not
mean the model made things up. We left the weights alone rather than silently redefine the unit every
historical number was measured in.must_not_contain
scan for 27 author-curated landmine phrases — plausible-but-false statements a
fabricating generator might emit. Not one was tripped anywhere. But 16 documents
with zero failures bounds the true rate only at ≤17.1%
(exact one-sided 95% Clopper–Pearson). That is not evidence of a general absence of fabrication.citable: false. The deterministic criteria, the cost ledger and the latency numbers are
unaffected.coverage is also a
completeness measure and saturates at 1.0 on the strong arms. So 0.50 of the weight sits on two
correlated completeness measures, and 0.10 on the only deterministic hallucination check.gpt-5.6-luna versus qwen3.7-flash differs by
+0.0003. Both intervals span zero. Decide on cost, speed or
operational risk, because quality is not deciding it for you.Every figure on this page is emitted by a build script from committed run scorecards, and a verifier recomputes each one from the raw judged cells before publication — including the statistics, which are derived from the per-document rows rather than copied from a stored aggregate. Prose and framing are hand-written and are not covered by that verifier, which is how the two errors listed in the correction note below got in.
Corrections. An earlier version of this page said the empty-artefact runs
"passed their own suite" — they did not, suite_passed is false on both — and
argued the top-two tie by comparing a 16-document mean difference against a per-document noise floor,
which is the wrong denominator. Both are fixed above, and the verifier now checks the first and computes
the second. Kept here because a page arguing that unverifiable numbers are not evidence does not get to
edit its own history quietly.
| Arms | 14 models, 16 documents each |
| Judge | anthropic:claude-sonnet-4-5-20250929 — cross-vendor, no arm judges itself |
| Judged criteria | faithfulness, coverage (0.5 of rubric weight) |
| Rubric weights | faithfulness 0.25, coverage 0.25, metadata_accuracy 0.2, anti_hallucination 0.1, format_compliance 0.1, citation_validity 0.05, token_efficiency 0.05 |
| Trials | 1 per document per arm |
| Metered spend | $6.7306 across all 14 arms (two are floors) |
| Judge calibration | never computed — kappa gate 0.6 unmet, citable: false |