The clause was one line long and we wrote it ourselves. The efficiency criterion for a specialist compared it against “the best-scoring measured peer” on its own axis, and nobody had asked what happens when two peers score the same. Then two did. gemma-2-2b, at 2.6 billion parameters, and Ministral-8B, at eight billion, both landed on exactly 34.03.
Against gemma, our candidates cleared the threshold by two to four times. Against Ministral, by about six, and everything passed.
The clause permitted either. That is the whole problem. A criterion that lets the person being judged choose the flattering reading after the numbers are in is not a criterion; it is a formality with a number attached. We took gemma — the tied peer that makes our ratio smallest — logged the ambiguity in the gate’s change log with its direction stated (harder to enter), and disclosed that the drafting was ours and had been resolved against our own interest.
A benchmark number is only as good as what was decided before it existed. The score is the least interesting part of a benchmark. Everything that makes it worth reading — the threshold, the dataset, the decoding, what happens to a model the harness cannot run, whether a third-party figure may sit in the same column as yours — was either fixed before the run or was available to be fixed afterwards, by someone with an interest in the outcome.
What has to be decided first
Our publication gate reads its criteria from a file with a declared_at timestamp, and it refuses criteria authored after the results they judge. There is one exception: results can be declared as pre-existing historical evidence, and evidence declared that way can only ever fail a gate, never pass one. That asymmetry is deliberate. Old results can convict; they cannot acquit.
The dataset is pinned by hash. Decoding is greedy, with the seed recorded, and identical across every arm. Every run records whether it was the full set, and a subset run never satisfies the gate. Subsets exist for iteration speed and for nothing else, because a 40-problem slice of HumanEval and the full 164 can disagree by more than the margin a criterion turns on. The public literature has a harsher version of the same point: harness version or a single separator character can move a result by one or two points, which is about the size of the gap between the top two rows on most boards.
When we ran the 3B capacity experiment, the criteria were written against the 3B base’s own measured HumanEval — 73.78, measured, not assumed — before the tuned run existed, and the config diff against the 1.5B run was checked to contain only the base model path. That is what it takes for a comparison between two runs to mean anything. It is tedious. It is also the entire difference between an experiment and an anecdote.
The row that is listed but not ranked
The Field Index ranks eleven models on three axes we run ourselves — code correctness, structured output, governed refusal — with weights declared before the aggregator ran. gemma-2-2b is not one of the eleven. It is listed underneath them, unranked, with two of its three axes filled in and a sentence explaining the third: its HumanEval run dies inside an mlx_lm batch-generation statistics path with a division-by-zero error. That is a library bug. It is not a capability result.
The tempting move is to score it zero and keep the table tidy. We call that the common-axis rule, and it runs the other way: a model is ranked only on axes measured for every ranked model, and a model missing one is shown as incomplete, never imputed, never back-filled, never zeroed for a harness failure. Zero would have been a false finding about a model that, on the axis we could measure, tied for best among the general peers. The same rule applied when gemma could not be measured on the routing suite because its chat template rejects a system role: recorded as not measurable, with the reason, not as 0.
There is a second kind of honesty the table has to carry. Four of the eleven ranked entries are our own fine-tunes, trained on data built for the structured-output axis, which carries 35% of the weight. They sit above the general models. That is expected and it is close to tautological, so the table says so above the first row, and points the reader at the only comparison that is evidence: a tuned entry against the base it was trained from. Cross-family comparisons on that table are context.
Citing is not aggregating
The largest public indices are mostly aggregates: a composite over ten or so evaluations, with weights the index operator chose, run on internal copies of datasets with no contamination testing disclosed. One benchmark-tracking site declines to list the best-known composite at all on the grounds that it is not a benchmark-native row. A February 2026 study from ETH Zürich and Stanford found 29 of 60 common benchmarks saturated: the gap between the top two models sits inside noise. On the repository-task front, one tracker lists 113 self-reported SWE-bench Verified scores and zero independently verified.
None of that makes those numbers useless. It makes them citations. A citation may sit beside our number with its source, date and link, visually distinct, and it may inform a shortlist. What it may not do is enter a composite we publish as ours, because we do not know the harness, the prompt, the decoding or the contamination controls behind it, and aggregating a number you cannot defend produces a number you cannot defend. The rule on our side is one line: only measurements we run ourselves enter the score.
That rule is why the table is small. Eleven ranked models is not a lot. It is the number we could run end-to-end on one machine under one harness, and every one of the rows is reproducible on a workstation with the pinned datasets.
The empty seat
On 13 March 2025 the Hugging Face Open LLM Leaderboard retired, after 13,000-plus models had been submitted to it. It was flawed in the ways every board is flawed, a fixed set of suites that models learned to fit, but it was the one place an open 2B or 7B was measured by someone other than its author, under one harness, with the results public. Eighteen months on, it has no single successor. The crowd-vote arena measures preference and has a documented problem with private-variant testing. The paid indices measure API endpoints, English text only, with no locally-run or quantized rows and no parameter or licence filters. The academic work on quantization quality, structured-output reliability and refusal drift exists as papers, not as a maintained table. Only one hub publishes per-question logs.
The seat is empty because the discipline that fills it is expensive and unglamorous — pinned hashes, full-set runs, a change log that records when a grader change moved a verdict in your favour — and because the people best placed to fill it usually have a model in the race.
The counter-argument is fair: a lab-only index of eleven models, three axes, two of them suites we authored, is a modest thing to set against boards with hundreds of entries. It is. Our authored suites measure what we chose to value, and that is stated wherever they are reported. But a reader who wants to know whether a 2B can hold a schema under adversarial input, and what it paid in coding ability to do so, cannot currently find that on any board with hundreds of entries. They can find it on ours, with the criteria dated before the runs, and they can re-run it.
The tie-break clause sits in the change log as a proposal, not yet in force. We hold it there for the same reason we will not settle the larger question of what a specialist owes on general capability while a candidate is pending: a criterion changed under pressure is a criterion changed under pressure, even when the change goes against us. It takes effect when nothing is waiting on it. The full method, including what a citation may and may not do on our pages, is on the methodology page.