14 models · 16 gold-labelled documents · cross-vendor judge · $6.73 of metered API spend

The LLM wiki got a spec
before it got a benchmark.

The pattern went from a viral post to a Google specification in ten weeks, and produced a pile of build-guides along the way. I went looking for a published comparison of which model to generate one with, what that costs, and how it fails, and did not find one — if you know of one I would genuinely like to read it. So this is the measurement I ran instead, including the part where the benchmark turns on itself.

What we ran

One document in, one structured page out: YAML frontmatter plus a markdown body, produced by a four-step loop (survey, extract, relate, compose). Fourteen candidate models generated the same 16 gold-labelled documents through the same loop and were scored by an identical judge prompt (sha256 12d7ddaa… on every arm) against a seven-criterion weighted rubric. The judge is anthropic:claude-sonnet-4-5-20250929 — a different vendor from every arm it scores, so no model grades its own homework. Cost and latency come from a per-call ledger, not from a price list.

Before any of the numbers below: that judge has never been calibrated against a human. No kappa exists, it carries 0.5 of the rubric weight, and every scorecard behind this page reports citable: false. What that does and does not invalidate is in What this does not tell you, and you should read it before you quote anything here.

The expensive model won by 0.019, and this instrument cannot tell you that is a win.

gpt-5.5 scores 0.9366 at $0.2100 per document. gpt-5.6-luna scores 0.9176 at $0.0147 — 14.3× cheaper, twice as fast, and not separated by this measurement. Paired across the same 16 documents the difference is +0.0190 with a 95% confidence interval of [-0.0100, +0.0480]. The interval spans zero. With one trial per document and an uncalibrated judge, that is not a narrow victory — it is a difference this instrument cannot resolve.

That is the practical finding: on this task, above a certain floor, the quality axis flattens and the cost axis does not. 8 of the fourteen arms have comparable means and land within 0.06 of the top score, and the spend across those 8 varies by 48×. If you picked your generator by reputation, you probably paid a multiple for a difference this benchmark cannot resolve.

Cheapest viable
$0.0044
per document, at 0.9173 — statistically level with the top arm
Cheapest to dearest
58×
on cost — but those two arms are only 0.137 apart in score. Cost and quality are barely related here
Fastest
11.5s
per document — the slowest arm takes 33× longer
Empty artefacts
16
files produced with zero content while the pipeline reported success

Quality against cost

Cost is a log scale — the spread is two orders of magnitude. Up and to the left is better. Models that failed to produce a usable page on every document are drawn hollow and faint: they are not viable at any price. Two arms crashed on documents the others completed, so their cost is a proven floor, not a total — those carry a right-pointing whisker, and their true position lies somewhere to the right of the marker.

0.95 0.86 0.77 0.68 0.59 0.50 $0.003 $0.01 $0.03 $0.10 $0.30 measured cost per document (log scale)rubric score gpt-5.5 gpt-5.6-luna ← 14× cheaper, tied qwen3.7-flash deepseek-v4-flash gpt-4.1 qwen3-32b gemini-3.1-flash-lite gemini-2.5-flash-lite gpt-5.4-nano seed-1.6-flash gpt-4.1-nano gpt-4o-mini gpt-5-nano kimi-k2.6
clears the document gate on all 16 — viable fails the gate — not viable at any price cost is a proven floor — true point lies right

All fourteen

Sorted by rubric score. Two rules apply to every row with no exceptions: the document gate is the number of documents scoring ≥ 0.75 divided by 16 — the whole corpus, never the survivors — and cost per document is total ledger spend divided by 16. An arm that cannot clear the gate consistently is not viable at any price, which is why the cheap end of this table is mostly unusable.

ModelScoreDoc gate$/docScore per $ LatencyOutput tokDocs scoredEmpty
gpt-5.5premium baseline0.9366100%$0.2100484.8s113,05516/16
gpt-5.6-lunatied with the top arm at 1/14th the cost0.9176100%$0.01476239.6s127,37016/16
qwen3.7-flashties luna, 3.3x cheaper, 2.6x slower · one OpenRouter route · rate triples above 32,000 input tokens0.9173100%$0.0044208102.1s248,99316/16
deepseek-v4-flashone OpenRouter route, DeepInfra fp4 — quantisation is a real variable here0.9102100%$0.0048190218.1s204,94716/16
gpt-4.1extract gpt-4.1-mini0.8979100%$0.06401468.4s99,03016/16
qwen3-32bthe only arm given more than one attempt — two earlier routes failed on 2026-08-10, this is the third; slowest arm0.889694%$0.0082108386.5s116,02816/16
gemini-3.1-flash-litefastest arm0.8816100%$0.01316711.5s55,95016/16
gemini-2.5-flash-lite2 crashes0.8798 †81%≥$0.006832.3s155,86314/16
gpt-5.4-nano0.876788%$0.01545762.2s127,29916/16
seed-1.6-flashJSON top-level was a string; 2 crashes · one OpenRouter route · output rate more than doubles above 128,000 input tokens0.8098 †44%≥$0.004154.2s103,65114/16
gpt-4.1-nanoretires 2026-10-230.799375%$0.003622224.2s37,45416/16
gpt-4o-miniwidely deployed default0.769269%$0.005314526.5s38,82716/16
gpt-5-nano5 of 16 artefacts came back EMPTY · scheduled shutdown 2026-12-110.674062%$0.008777120.3s294,75316/165 EMPTY
kimi-k2.611 of 16 artefacts came back EMPTY · one OpenRouter route; rate includes a 39% promotional discount, non-promotional routes run ~39% higher on output0.525931%$0.05759298.4s325,66716/1611 EMPTY

† This arm's score is a mean over a self-selected subset. It crashed on documents the other arms completed, so the surviving documents are an easier corpus than the one everybody else was measured on. These means are not comparable to the rest of the table in either direction — the crashes may have removed the hardest documents or the easiest, and nothing here distinguishes the two.

≥ The run crashed mid-document, so a total cost does not exist. The figure shown is what the ledger proves was already billed, divided by all 16 documents — a floor, never a total. Score per $ is — rather than a number, because a ratio built on a floor would read as a measurement.

Output tok counts only documents that completed. Crashed documents burned tokens too — a further 3,049 for gemini-2.5-flash-lite and 1,174 for seed-1.6-flash — and produced nothing.

Three things the price list cannot tell you

Every one of these was invisible until the models actually ran. None of them is derivable from a rate card, a context-window number, or a public leaderboard.

Cheap can mean lazy

gpt-4o-mini emitted 38,827 output tokens across the corpus and gpt-4.1-nano emitted 37,454, against gpt-5.6-luna's 127,370 on the identical documents — roughly a third. They are cheaper partly because the artefact is thinner, and they miss the document gate 31% and 25% of the time. Cost per token is not cost per usable page.

Reasoning tokens break the rate card

gpt-5-nano lists at a fraction of the recommended model's rate and emitted 294,753 output tokens — 2.3× as many as gpt-5.6-luna, all billed at the output rate. It ended up the lowest-scoring OpenAI arm at 0.6740 and among the slowest at 120.3s per document. A published price is an arithmetic input, not a forecast.

A big window is not structured output

gemini-2.5-flash-lite advertises a million-token window and failed on two of sixteen documents — for two different reasons. One was size: JSON truncated at character 211,025 on the largest document. The other was not: on a roughly 1,000-token blog post it returned well-formed JSON using a claim_type value outside our schema's enum, and our validator rejected it. Call that the model's fault or our schema's — window size explains neither.

Two models were priced and live, and still would not return an object

Both appear in the public price book. Both accept requests. Neither returned a structured artefact, which is the entire job. Both refused models were probed through one broker, in one mode, on 2026-08-07. Neither was tested against its vendor's own API, so this is a statement about those routes on that date and not about the models in general.

ling-3.0-flash — the cheapest row in the book

'does not support feature: structured-outputs'. response_format is absent from supported_parameters on both its endpoints.

qwen3.5-flash — answers, but not with an object

In JSON mode it returns the bare float -1.0000000000000002e+308 instead of an object. Five attempts out of five, two system prompts, two user messages, finish_reason: "stop", billed in full. A single endpoint, so there is no other route to try.

The rule these produced, in order: a rate row proves arithmetic, not availability — then availability does not prove usability — then an endpoint that answers is not an endpoint that works.

The failure mode a green build will not catch

This is the one that would have shipped, and it is the reason a benchmark has to check the artefact and not just the exit code.

16 artefacts came back with an empty body — and every single row said ok.

kimi-k2.6 produced 11 of 16. gpt-5-nano — from a major lab, not an exotic endpoint — produced 5 of 16. These are not zero-byte files, which is what makes them dangerous: they carry valid YAML frontmatter and score 1.0 on schema validity, 1.0 on citation validity and up to 0.89 on metadata accuracy, with init_md_token_count: 0 and nothing whatsoever in the body. A downstream indexer accepts them precisely because the envelope is well-formed.

Each burned 19,592 to 30,364 output tokens of output. So the failure is not merely silent, it is expensive and silent: real money spent, an empty-bodied artefact indexed, and no error, no warning and no non-zero exit to say so. The only trace is a token count of zero on a field nothing was checking.

One layer did catch it, and it is worth being precise about which. The row status lied and the error count lied, but the aggregate document gate held — suite_passed: false on both arms — and only because this benchmark happens to score the artefact itself. A pipeline checking exit codes and token counts would have shipped all 16. The fix belongs in every generator of this kind: assert the artefact is non-empty before recording the step as successful, in the worker and in the benchmark, because a benchmark-only fix leaves the product broken.

Both arms ran one configuration, one trial per document, on 2026-08-07: gpt-5-nano on its listed deployment, kimi-k2.6 through a single OpenRouter route. This benchmark cannot separate model behaviour from route or broker behaviour, and neither arm was re-tested on an alternate endpoint.

What this does not tell you

The section that earns the rest of the page. Everything below is a limit we found in our own instrument, and the first item cost us a headline number.

Reproduction

Every figure on this page is emitted by a build script from committed run scorecards, and a verifier recomputes each one from the raw judged cells before publication — including the statistics, which are derived from the per-document rows rather than copied from a stored aggregate. Prose and framing are hand-written and are not covered by that verifier, which is how the two errors listed in the correction note below got in.

Corrections. An earlier version of this page said the empty-artefact runs "passed their own suite" — they did not, suite_passed is false on both — and argued the top-two tie by comparing a 16-document mean difference against a per-document noise floor, which is the wrong denominator. Both are fixed above, and the verifier now checks the first and computes the second. Kept here because a page arguing that unverifiable numbers are not evidence does not get to edit its own history quietly.

Arms14 models, 16 documents each
Judgeanthropic:claude-sonnet-4-5-20250929 — cross-vendor, no arm judges itself
Judged criteriafaithfulness, coverage (0.5 of rubric weight)
Rubric weightsfaithfulness 0.25, coverage 0.25, metadata_accuracy 0.2, anti_hallucination 0.1, format_compliance 0.1, citation_validity 0.05, token_efficiency 0.05
Trials1 per document per arm
Metered spend$6.7306 across all 14 arms (two are floors)
Judge calibrationnever computed — kappa gate 0.6 unmet, citable: false