Learn · Benchmarking

HumanEval 57.93 to 3.05. The gate refused, and we kept the receipt.

Our own publication gate turned down our own model, the decision memo told the founder no, and a rename broke a signature chain. Why a lab that prints its failing numbers is the one whose passing numbers are worth reading.

LeanLogix6 min read

The candidate was good. On the axis we would actually sell — turning a hostile, messy transcript into an exact six-key JSON disposition without being steered — a 1.5B model we had trained through our own governed loop scored 82.67 where Ministral-8B scored 30.97 and Llama-3.1-8B 28.41. It was the only model in the comparison that obeyed zero smuggled instructions; every peer obeyed five to eight. Measured by us, one harness, one prompt, peers a buyer could re-run in minutes.

Then the gate printed the rest of the verdict.

C2 general hold (humaneval): tuned 3.05 vs base 57.93 (drop 54.88; tolerance 3.0). HOLD. Publication refused. The adapter had learned “always emit the six-key JSON” so thoroughly that it emitted almost nothing else, and the first reviewer to try it off-script would have found that in minutes. The decision memo that went to the founder the same day said, in its first line, no — not yet. Two things had to land first, and neither was a matter of framing.

Publishing your own model’s failing number is the cheapest credibility you will ever buy. Not because humility is charming. Because every number on every board is reported by someone with an interest in it, and the only thing that makes a lab’s passing numbers believable is that it has been seen, in public, printing the failing ones under the same criteria.

What the gate is, and why it cannot be talked to

Our publication gate is a lifecycle stage, not a mood. It sits after certification and before any public distribution, and it reads its criteria from a file whose declared_at timestamp must precede the results it judges. The criteria are five: beat your own base on your specialised public axis by at least five points; hold within three points of your base on the general suite for your class, full set, plain prompt; be defensible against at least two open-weight peers in your size class that we measured ourselves; report the number your shipped posture actually produces; and be certified and signed. Any miss is HOLD, with the failing criteria named. There is no --force.

The founder’s bar, written into the document, is a single sentence: if we are publishing something it has to compare very well, and if we are still iterating we hold. The gate turns that into arithmetic so that nobody has to be brave on a Friday.

Since that first verdict, five trained artifacts have gone through it and five have failed the same criterion. The 1.5B intake line dropped 54.88, then 57.93, then 17.08 points below its base across three cycles. A 3B run built to test whether size was the problem dropped 20.12. Bearing, our router, already signed and in service before the framework existed, dropped all 57.93; it emits nothing at all on coding prompts. Every one of those verdicts is in the repository with its numbers, and so is the sentence that makes them uncomfortable: at a tolerance of 17.1 points the third cycle would have passed, and at about 20.2 the 3B would have. We wrote that down before anyone could be accused of discovering it later, and we declined to change the tolerance while a candidate was pending, because a framework decision made with a known beneficiary is not a framework decision.

The rename that broke a signature

The most instructive failure was not a score. When we re-released our portfolio router under the name Bearing — a naming exercise, not a training run, the cheapest real win on the table — it failed the lifecycle criterion as unsigned. The weights were identical to the signed release. The signature did not care. It is bound to the release’s model_id and its passport digest, and a passport with a different name has a different digest, so the same bytes under a new label are legitimately unsigned.

That is the signature working. A signature that survived renaming would be a signature that survived relabelling a weaker model as a stronger one. But it means that each of the nine further renames planned for the series is a fresh platform signature and a human approval — a person who is not the trainer, recorded in the lineage — and not a text edit. We would rather have learned that from our own gate than from an auditor’s.

A header that is printed before the table

The same instinct governs how the Field Index is allowed to describe itself. The index is designed to combine benchmarks we run with signals from real enterprise delivery: how often a model’s changes are accepted unmodified, how many review rounds, whether it holds its operating contract. That field layer is defined in the specification, gated behind a narrower consent scope than training, computed only over cohorts of at least three client organisations and two hundred tasks, and clamped to fifteen points either way so that private, unreproducible signal can inform a ranking without swamping the reproducible part.

It is also, today, empty.

So before the first row of the table there is a coverage header, and it reads: 0 of 11 entries field-weighted (0.0%) · 11 lab-only. Field cohort: field layer not populated. Every published edition carries one, and when it reads zero the index is published as the lab-only edition and the field-weighting is described as defined but not yet populated. The rule behind it is short: the name may not carry an implication the coverage line contradicts. A per-row label saying lab-only is not sufficient when every row is lab-only, because the name still implies data we do not have.

The clamp has a disclosure rule of its own. When a model’s computed field adjustment exceeds the bound, the entry does not quietly show the bound. It shows Δ = +15 (clamped from +21.4), with a footnote, because a model whose real-world behaviour diverges that sharply from its lab score is the most interesting row in the table and hiding the divergence inside a bound would suppress the one finding the index exists to surface.

One more receipt. When our first throughput sweep turned out to have run concurrently with another job on the GPU, two background chains released on the same process check, the base model had been understated by a third, and two claims built on it were wrong. They were withdrawn by name. The contended files were kept under an INVALIDATED.json marker rather than deleted, because the discrepancy between the contended run and the clean one is itself evidence of how much contention costs, and every measurement since records which GPU processes were present before and after it ran.

Why this is cheap, and who should care

The objection is that all of this looks like a lab talking itself out of shipping. Five verdicts, five holds, a signed model that cannot be renamed, an index that announces its own emptiness. A competitor with looser standards publishes the 82.67 and moves on.

They do, and it costs them the thing they were trying to buy. Consider the state of the boards. One tracker lists 113 self-reported scores on the best-known repository-task benchmark and zero independently verified. The one place open small models were measured by someone other than their authors retired in March 2025. Every published number in this field is a claim by an interested party, and buyers, auditors and the more sceptical kind of engineer have adjusted accordingly: a passing score from a lab with no visible record of failing ones is discounted heavily, whatever it says. The lab that publishes its holds is making a different offer. It is saying: here are the criteria, dated before the run; here is what we measured; here is the verdict that went against us and the tolerance that would have flipped it. Now read our passing number.

That offer is cheap to make. It costs a few paragraphs and the occasional argument with oneself. And it is the offer an investor is actually being asked to underwrite — not a model, which will be beaten by a better one within the year, but a measurement practice that becomes more valuable each time it is exercised and that a competitor cannot copy by publishing a higher number. The 82.67 was never the asset. The gate that refused it was.

The candidate is still behind the gate. The benchmark it led is being generalised into a structured-output axis on the Field Index, where it measures every model in the table instead of justifying one of ours. When a model of ours passes, it will be listed with the same criteria, the same change log, and the same coverage line as everything that did not.

Evidence

Every figure in this guide is taken from the repository documents below, measured by us on one harness with identical prompts, greedy decoding, and pinned datasets, or from the external sources listed with them.

On the platform

Keep reading

Choosing · 6 minWe declared an order of magnitude. The measurement came back 2.16×.Training · 7 minCycle three, and the failure moved again.
All guides