14 models · 16 gold-labelled documents · cross-vendor judge · $6.69 of metered API spend

The LLM wiki got a spec
before it got a benchmark.

The pattern went from a viral post to a Google specification in ten weeks, and produced a pile of build-guides along the way. I went looking for a published comparison of which model to generate one with, what that costs, and how it fails, and did not find one — if you know of one I would genuinely like to read it. So this is the measurement I ran instead, including the part where the benchmark turns on itself.

What we ran

One document in, one structured page out: YAML frontmatter plus a markdown body, produced by a four-step loop (survey, extract, relate, compose). Fourteen candidate models generated the same 16 gold-labelled documents through the same loop and were scored by an identical judge prompt (sha256 12d7ddaa… on every arm) against a seven-criterion weighted rubric. The judge is anthropic:claude-sonnet-4-5-20250929 — a different vendor from every arm it scores, so no model grades its own homework. Cost and latency come from a per-call ledger, not from a price list.

Before any of the numbers below: that judge has never been calibrated against a human. No kappa exists, it carries 0.5 of the rubric weight, and every scorecard behind this page reports citable: false. What that does and does not invalidate is in What this does not tell you, and you should read it before you quote anything here.

The expensive model won by 0.012, and this instrument cannot tell you that is a win.

gpt-5.5 scores 0.9073 at $0.2152 per document. gpt-5.6-luna scores 0.8953 at $0.0153 — 14.1× cheaper, 1.4× faster, and not separated by this measurement. Paired across the same 16 documents the difference is +0.0120 with a 95% confidence interval of [-0.0353, +0.0593]. The interval spans zero. With one trial per document and an uncalibrated judge, that is not a narrow victory — it is a difference this instrument cannot resolve.

That is the practical finding: on this task, above a certain floor, the quality axis flattens and the cost axis does not. 6 of the fourteen arms have comparable means and land within 0.06 of the top score, and the spend across those 6 varies by 44×. If you picked your generator by reputation, you probably paid a multiple for a difference this benchmark cannot resolve.

Cheapest viable
$0.0049
per document, at 0.8740 — statistically level with the top arm
Cheapest to dearest
58×
on cost — but those two arms are only 0.113 apart in score. Cost and quality are barely related here
Fastest
12.0s
per document — the slowest arm takes 26× longer
Empty artefacts
19
generations that returned a valid envelope with nothing in the body — now rejected before scoring, and scored as successes in the previous run

Quality against cost

Cost is a log scale — the spread is two orders of magnitude. Up and to the left is better. Models that failed to produce a usable page on every document are drawn hollow and faint: they are not viable at any price. 7 arms crashed on documents the others completed, so their cost is a proven floor, not a total — those carry a right-pointing whisker, and their true position lies somewhere to the right of the marker.

0.95 0.86 0.77 0.68 0.59 0.50 $0.003 $0.01 $0.03 $0.10 $0.30 measured cost per document (log scale)rubric score gpt-5.5 gpt-5.6-luna <-- 14x cheaper, tied qwen3.7-flash gpt-4.1 qwen3-32b gemini-3.1-flash-lite gemini-2.5-flash-lite gpt-5.4-nano seed-1.6-flash gpt-4.1-nano gpt-4o-mini gpt-5-nano kimi-k2.6
clears the document gate on all 16 — viable fails the gate — not viable at any price cost is a proven floor — true point lies right

All fourteen

Sorted by rubric score. Two rules apply to every row with no exceptions: the document gate is the number of documents scoring ≥ 0.75 divided by 16 — the whole corpus, never the survivors — and cost per document is total ledger spend divided by 16. An arm that cannot clear the gate consistently is not viable at any price, which is why the cheap end of this table is mostly unusable.

ModelScoreDoc gate$/docScore per $ LatencyOutput tokDocs scoredEmpty
gpt-5.5premium baseline0.9073100%$0.2152471.7s117,26116/16
gpt-4.1extract gpt-4.1-mini0.9066100%$0.06411460.8s98,36916/16
kimi-k2.611 of 16 generations returned an empty body · scored on 5 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · one OpenRouter route; rate includes a 39% promotional discount, non-promotional routes run ~39% higher on output0.9023 †25%≥$0.0554206.2s56,5565/1611 EMPTY
gpt-5.4-nano0.9010100%$0.01545952.9s126,73116/16
gpt-5.6-lunatied with the top arm at 1/14th the cost0.895394%$0.01535952.6s133,30016/16
qwen3-32b3 failed on malformed output or schema validation · scored on 13 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · the only arm given more than one attempt — two earlier routes failed on 2026-08-10, this is the third; slowest arm0.8745 †81%≥$0.0063306.7s79,16213/16
qwen3.7-flashone OpenRouter route · rate triples above 32,000 input tokens0.874081%$0.004917891.8s282,76616/16
gpt-5-nano7 of 16 generations returned an empty body · scored on 9 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · scheduled shutdown 2026-12-110.8708 †56%≥$0.008491.0s128,4139/167 EMPTY
gemini-3.1-flash-litefastest arm0.868394%$0.01336512.0s57,37416/16
seed-1.6-flash4 failed on malformed output or schema validation · scored on 12 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · one OpenRouter route · output rate more than doubles above 128,000 input tokens0.8505 †50%≥$0.004369.0s112,76012/16
gemini-2.5-flash-lite4 failed on malformed output or schema validation · scored on 12 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them0.8423 †62%≥$0.006134.8s143,87112/16
gpt-4.1-nano1 of 16 generations returned an empty body · 1 failed on malformed output or schema validation · scored on 14 of 16 documents, so its mean describes an easier corpus than the arms that completed all of them · retires 2026-10-230.7947 †62%≥$0.003722.9s33,88514/161 EMPTY
gpt-4o-miniwidely deployed default0.776862%$0.005514131.6s43,12716/16
deepseek-v4-flashNO MEASUREMENT - every document failed at the provider (HTTP 404, no route available for this model id) · one OpenRouter route, DeepInfra fp4 — quantisation is a real variable here0.0000 †0%≥$0.00000.0s00/16

† This arm's score is a mean over a self-selected subset. It crashed on documents the other arms completed, so the surviving documents are an easier corpus than the one everybody else was measured on. These means are not comparable to the rest of the table in either direction — the crashes may have removed the hardest documents or the easiest, and nothing here distinguishes the two.

≥ The run crashed mid-document, so a total cost does not exist. The figure shown is what the ledger proves was already billed, divided by all 16 documents — a floor, never a total. Score per $ is — rather than a number, because a ratio built on a floor would read as a measurement.

Output tok counts only documents that completed. Crashed documents burned tokens too — a further 3,049 for gemini-2.5-flash-lite and 1,174 for seed-1.6-flash — and produced nothing.

Three things the price list cannot tell you

Every one of these was invisible until the models actually ran. None of them is derivable from a rate card, a context-window number, or a public leaderboard.

Cheap can mean lazy

gpt-4o-mini emitted 43,127 output tokens across the corpus and gpt-4.1-nano emitted 33,885, against gpt-5.6-luna's 133,300 on the identical documents — roughly a third. They are cheaper partly because the artefact is thinner, and they miss the document gate 38% and 38% of the time. Cost per token is not cost per usable page.

Reasoning tokens break the rate card

gpt-5-nano lists at a fraction of the recommended model's rate and emitted 128,413 output tokens — 1.0× as many as gpt-5.6-luna, all billed at the output rate. It ended up the lowest-scoring OpenAI arm at 0.8708 and among the slowest at 91.0s per document. A published price is an arithmetic input, not a forecast.

A big window is not structured output

gemini-2.5-flash-lite advertises a million-token window and failed on two of sixteen documents — for two different reasons. One was size: JSON truncated at character 211,025 on the largest document. The other was not: on a roughly 1,000-token blog post it returned well-formed JSON using a claim_type value outside our schema's enum, and our validator rejected it. Call that the model's fault or our schema's — window size explains neither.

Two models were priced and live, and still would not return an object

Both appear in the public price book. Both accept requests. Neither returned a structured artefact, which is the entire job. Both refused models were probed through one broker, in one mode, on 2026-08-07. Neither was tested against its vendor's own API, so this is a statement about those routes on that date and not about the models in general.

ling-3.0-flash — the cheapest row in the book

'does not support feature: structured-outputs'. response_format is absent from supported_parameters on both its endpoints.

qwen3.5-flash — answers, but not with an object

In JSON mode it returns the bare float -1.0000000000000002e+308 instead of an object. Five attempts out of five, two system prompts, two user messages, finish_reason: "stop", billed in full. A single endpoint, so there is no other route to try.

The rule these produced, in order: a rate row proves arithmetic, not availability — then availability does not prove usability — then an endpoint that answers is not an endpoint that works.

The failure mode a green build will not catch

This is the one that would have shipped, and it is the reason a benchmark has to check the artefact and not just the exit code.

An artefact can be valid, billed, indexed — and empty.

19 generations returned a well-formed envelope with an EMPTY body (kimi-k2.6 11, gpt-5-nano 7, gpt-4.1-nano 1), across 11 distinct documents. They are not zero-byte files - the frontmatter parses, which is exactly why a downstream indexer would accept them. Between them they burned 417,388 output tokens (7,004 to 28,714 each) producing nothing. In the previous run these were scored as SUCCESSFUL documents; a guard now rejects an empty body before it can be scored, which is why they appear here as errors and why the affected arms report a lower completed-document count rather than a quietly lower score.

The previous version of this benchmark scored every one of them as a successful document. The row said ok, the error count said zero, and the only trace was a token count of zero on a field nothing was checking. A pipeline watching exit codes would have indexed all 19 and billed for them.

The fix is a guard that rejects an empty body before it can be scored, and it belongs in the worker as well as the benchmark — a benchmark-only fix leaves the product broken. That is why these appear here as errors rather than as quietly low scores, and why the affected arms report fewer completed documents instead of a slightly worse mean. Read the gate column, not the mean: an arm that fails a document is not scored on it, so failing more can look like scoring higher.

What this does not tell you

The section that earns the rest of the page. Everything below is a limit we found in our own instrument, and the first item cost us a headline number.

Reproduction

Every figure on this page is emitted by a build script from committed run scorecards, and a verifier recomputes each one from the raw judged cells before publication — including the statistics, which are derived from the per-document rows rather than copied from a stored aggregate.

Prose is the weak point and it is worth being explicit about it. A verifier that checks numbers cannot check a sentence, and an earlier draft of this page carried a whole narrative from a previous run underneath a correct table. The generator now rebuilds any sentence that restates a figure, and refuses to write a ledger where a superseded number survives in prose — but a claim whose premise has quietly stopped being true is still something only a reader can catch.

Supersession. This publication SUPERSEDES an earlier one measured on a different corpus. The earlier numbers do not carry over and should not be compared to these. What changed:

Said plainly because a page arguing that unverifiable numbers are not evidence does not get to edit its own history quietly. If you saw the earlier figures, they were measured on a corpus that included two documents this repository does not contain, which is why they are not reproducible here and are not repeated.

Arms14 models, 16 documents each
Judgeanthropic:claude-sonnet-4-5-20250929 — cross-vendor, no arm judges itself
Judged criteriafaithfulness, coverage (0.5 of rubric weight)
Rubric weightsfaithfulness 0.25, coverage 0.25, metadata_accuracy 0.2, anti_hallucination 0.1, format_compliance 0.1, citation_validity 0.05, token_efficiency 0.05
Trials1 per document per arm
Metered spend$6.6882 across all 14 arms (two are floors)
Judge calibrationnever computed — kappa gate 0.6 unmet, citable: false