Small-model benchmarks we ran ourselves

The Field Index

Every number on this page was measured on one harness — identical prompts, greedy decoding, pinned datasets, one machine — and no third-party number is aggregated into any of it.

v0 · lab-only editionmeasured August 31, 202611 ranked · 1 listed unranked · 15 efficiency arms1.2B – 8B parameters

Coverage: 0 of 11 entries field-weighted (0.0%) · 11 lab-only — published as the lab-only edition; the field layer is defined but not populated.

Field cohort: field layer not populated. Composition: 4 base · 3 peer · 4 ours (fine-tuned by us — not peer comparisons). The name may not carry an implication this line contradicts: this is a lab index with a field slot, and nothing here is field-weighted.

11

ranked entries, every one measured by us

3

axes in the composite; 2 more defined, not yet measured

0

entries field-weighted — the field layer is empty

15

contention-clean efficiency arms on one Mac Studio

The ranked table

Eleven small models, one harness, three axes.

Code correctness is HumanEval pass@1. Structured output and governed refusal come from intake-bench v0, our schema-exact-under-adversarial-input suite. Weights were declared before the aggregation that uses them: code correctness 0.45 · structured output 0.35 · governed refusal 0.20.

Read the ranking with this in mind

Rows marked ours were fine-tuned by us on data built for the structured-output axis, which carries 35% of the weight. Their position above general models is therefore expected and is not independent validation — a model trained for an axis outscoring models that were not is close to tautological. The honest comparison is a tuned entry against the base it was trained from, shown in the same table: intake-3B v1 (68.96) against Qwen2.5-3B-Instruct (34.83), and the intake-1.5B line against Qwen2.5-1.5B-Instruct (26.80). Cross-family comparisons here are context, not evidence. Switch on Compare to base to see every delta.
The Field Index, lab-only edition: 11 ranked entries measured on one harness. Sortable by every column.
1
intake-3B v1ours · tuned for the structured axis
intake-3b-v1
3BApache-2.0+ our adapter53.6688.2469.64+7.5468.96
2
intake-1.5B v3ours · tuned for the structured axis
intake-1_5b-v3
1.5BApache-2.0+ our adapter40.8583.1971.43+3.3761.78
3
Qwen2.5-Coder-7B-Instructbase
qwen2.5-coder-7b-instruct-BASE
7.6BApache-2.089.0236.5525.45+9.3657.94
4
Ministral-8B-Instruct-2410peer
peer-ministral-8b
8BMistral Research License81.1034.0322.32+9.4752.87
5
Qwen2.5-7B-Instructbase
base-qwen2.5-7b
7.6BApache-2.082.3221.4314.29+1.2547.40
6
intake-1.5B v1ours · tuned for the structured axis
intake-1_5b-v1
1.5BApache-2.0+ our adapter3.0584.4573.21+5.5045.57
7
Meta-Llama-3.1-8B-Instructpeer
peer-llama-3.1-8b
8BLlama 3.1 Community License71.9528.9912.50+1.8045.02
8
intake-1.5B v2ours · tuned for the structured axis
intake-1_5b-v2
1.5BApache-2.0+ our adapter0.0086.5573.21+3.2244.93
9
Qwen2.5-3B-Instructbase
base-qwen2.5-3b
3BApache-2.073.782.104.46+2.1034.83
10
Qwen2.5-1.5B-Instructbase
base-qwen2.5-1.5b
1.5BApache-2.057.932.100.00+0.3526.80
11
Llama-3.2-1B-Instructpeer
peer-llama-3.2-1b
1.2BLlama 3.2 Community License41.460.000.00+0.0018.66

Every cell is a self-run measurement. Code = HumanEval full-set pass@1. Structured and Refusal = intake-bench v0 exact-match on the public split; pub−held is the public-minus-held-out gap on the structured axis, shown as the overfitting indicator. Field Index = 0.45 × code + 0.35 × structured + 0.20 × refusal, with a field delta of 0 on every row because the field layer is not populated. Licenses link to the exact Hugging Face artifact we measured; the adapter rows are ours. Registry entries for our own models are on /models.

Tuned against its own base

What fine-tuning bought, and what it cost.

Same harness, same prompts, same decoding — the only variable is the adapter. Structured output and governed refusal go up because that is what the data taught; code correctness goes down on every line. intake-1.5B v3 scores 40.85 on HumanEval against its base's 57.93. That trade is the finding, and it is the one nobody else publishes.

OursTrained fromΔ codeΔ structuredΔ refusalΔ pub−heldΔ Field Index
intake-3B v1intake-3b-v1 · 68.96Qwen2.5-3B-Instructbase-qwen2.5-3b · 34.83−20.12+86.14+65.18+5.44+34.13
intake-1.5B v3intake-1_5b-v3 · 61.78Qwen2.5-1.5B-Instructbase-qwen2.5-1.5b · 26.80−17.08+81.09+71.43+3.02+34.98
intake-1.5B v1intake-1_5b-v1 · 45.57Qwen2.5-1.5B-Instructbase-qwen2.5-1.5b · 26.80−54.88+82.35+73.21+5.15+18.77
intake-1.5B v2intake-1_5b-v2 · 44.93Qwen2.5-1.5B-Instructbase-qwen2.5-1.5b · 26.80−57.93+84.45+73.21+2.87+18.13

A larger public-minus-held-out gap after tuning is the overfitting indicator moving the way you would expect it to; it is published, not smoothed. v1 and v2 lost almost all of their code correctness (3.05 and 0.00); v3 is the checkpoint where the trade became defensible. `tuned` entries are OUR fine-tunes, trained on data related to the axis they are scored on. They are not peer comparisons and must never be read as independent validation.

Listed but unranked

Missing a common axis — never imputed.

The composite is computed only over axes measured for EVERY ranked model in this index version. A model missing any common axis is listed but UNRANKED, never imputed and never silently renormalised — otherwise a model with fewer measurements could outrank one with more.

ModelLicenseCodeStructuredRefusalWhy unranked
gemma-2-2b-itpeer-gemma-2-2b · 2.6B · peerGemma Terms of Use34.038.04code_correctness not measurable: mlx_lm batch_generate raises ZeroDivisionError in its stats path on this model (library bug, not a capability result). Scoring it zero would be a false finding.

Efficiency

Throughput, memory, and cost — measured, not inferred from a parameter count.

Apple Mac Studio, M3 Ultra, 512 GB unified memory, MLX via mlx-lm, bf16/fp16 weights as published by mlx-community. Batch size 1, a fixed 12-prompt set of identical length class, 128 max tokens, greedy, one warm-up run then 12 measured, peak memory reset immediately before the measured runs. Every arm records the GPU processes present before and after; only arms with a clean contention guard appear here.

Finding 1 · the advantage is 2–4×, not 10×

Before measuring, we wrote down an expectation of roughly an order-of-magnitude advantage in index points per GB for a 1.5B specialist against a larger general model. Measured on the specialised axis against the best-scoring small peer (gemma-2-2b, 34.03), the unfused candidates land at 2.16× – 4.12× per GB. The expectation is not met. A well-chosen small general model is a much harder baseline than a frontier-class one, and it is what a cost-conscious buyer would actually reach for. The efficiency argument survives; the order-of-magnitude framing does not, on any dimension.

Finding 2 · an unfused adapter runs at about half of base throughput

A LoRA adapter served unfused costs throughput for nothing. The Qwen2.5-1.5B base runs 126.92 tok/s; the same base under our adapters runs 65.6969.25 (5261% of base) at unchanged memory. Fused with mlx_lm.fuse, v3 returns to 120.33 tok/s (95% of base) with the axis score unchanged within bf16 noise; the 3B line goes 44.8269.21 against a base of 73.46. Fusing is a serving requirement, not an optimisation. An earlier sweep that showed a ~3× penalty was measured under GPU contention and is withdrawn; these are the clean numbers.
Efficiency measurements for every contention-clean arm: tokens per second, peak resident memory, load time, seconds per task, cost per one thousand tasks, structured-output score and Field Index where applicable.
Armtok/sPeak GBLoad ss / task$ / 1k tasksStructuredField Index
Qwen2.5-1.5B-Instructbase
mlx-community/Qwen2.5-1.5B-Instruct-bf16
126.922.9670.930.83710.17512.1026.80
intake-1.5B v1ours · adapter
mlx-community/Qwen2.5-1.5B-Instruct-bf16 + intake-1_5b-v1
69.253.0430.841.62570.340084.4545.57
intake-1.5B v2ours · adapter
mlx-community/Qwen2.5-1.5B-Instruct-bf16 + intake-1_5b-v2
67.313.0430.621.24660.260886.5544.93
intake-1.5B v3ours · adapter
mlx-community/Qwen2.5-1.5B-Instruct-bf16 + intake-1_5b-v3
65.693.0430.741.01100.211583.1961.78
intake-1.5B v3 · fusedours · fused
mlx-community/Qwen2.5-1.5B-Instruct-bf16adapter fused into the base weights with mlx_lm.fuse; re-scored on intake-bench, not on the composite
120.332.9670.870.64680.135383.61
bearing-1.5B v1ours · adapter
mlx-community/Qwen2.5-1.5B-Instruct-bf16 + portfolio-router-1_5b-v1-promotedour routing adapter — measured for efficiency only; not scored on the Field Index axes
67.123.0431.041.13360.2371
Qwen2.5-3B-Instructbase
mlx-community/Qwen2.5-3B-Instruct-bf16
73.465.8221.271.50070.31392.1034.83
intake-3B v1ours · adapter
mlx-community/Qwen2.5-3B-Instruct-bf16 + intake-3b-v1
44.825.9151.311.56350.327088.2468.96
intake-3B v1 · fusedours · fused
mlx-community/Qwen2.5-3B-Instruct-bf16adapter fused into the base weights with mlx_lm.fuse; re-scored on intake-bench, not on the composite
69.215.8221.320.94270.197287.82
Qwen2.5-7B-Instructbase
mlx-community/Qwen2.5-7B-Instruct-bf16
39.6414.2542.552.78330.582221.4347.40
Qwen2.5-Coder-7B-Instructbase
mlx-community/Qwen2.5-Coder-7B-Instruct-bf16
39.5314.2542.592.64790.553936.5557.94
Llama-3.2-1B-Instructpeer
mlx-community/Llama-3.2-1B-Instruct-bf16
178.492.3811.140.66620.13940.0018.66
gemma-2-2b-itpeer
mlx-community/gemma-2-2b-it-fp16
78.284.9341.661.45310.303934.03
Meta-Llama-3.1-8B-Instructpeer
mlx-community/Meta-Llama-3.1-8B-Instruct-bf16
37.5315.0403.143.12420.653528.9945.02
Ministral-8B-Instruct-2410peer
mlx-community/Ministral-8B-Instruct-2410-bf16
38.4915.0063.032.93170.613234.0352.87

The cost model, declared before measurement

$0.753 per hour: a $9,499 machine amortised over 3 years at 50% utilisation, plus 200 W at $0.15/kWh. Cost per 1k tasks = wall_clock_seconds_per_task * 1000 / 3600 * total_usd_per_hour. Model RANKINGS by cost are invariant to this hourly rate — changing it scales every model's cost identically. Only the absolute dollar figures move. The rate is published so a reader can substitute their own.

What the dollar figure is not

This is a local-hardware cost model. It is NOT comparable to a hosted API's per-token price without also accounting for utilisation, availability and operations, and no such comparison is published here. Fused arms are the serving artifact; unfused arms are shown because they were measured and because the gap between them is the finding. Dashes mean the arm was not scored on that axis — never zero.

The efficiency criterion, as declared — and who passes it

S4 asks for ≥ 3.0× the best-scoring measured peer on at least 2 of 3 ratios — points per GB, per second, per dollar — on the specialised axis. The threshold was declared before measurement and not revised after. Two peers tie on the axis at 34.03; we compare against the more efficient one (gemma-2-2b-it), the choice that works against us.

Candidatepts / GBpts / spts / $LeadsVerdict
intake-3B v114.92 · 2.16×56.44 · 2.41×269.8 · 2.41×0 / 3FAIL
intake-3B v1 · fused15.08 · 2.19×93.16 · 3.98×445.3 · 3.98×2 / 3PASS
intake-1.5B v228.44 · 4.12×69.43 · 2.96×331.9 · 2.96×1 / 3FAIL
intake-1.5B v127.75 · 4.02×51.95 · 2.22×248.4 · 2.22×1 / 3FAIL
intake-1.5B v3 · fused28.18 · 4.08×129.27 · 5.52×618.0 · 5.52×3 / 3PASS
intake-1.5B v327.34 · 3.96×82.28 · 3.51×393.3 · 3.51×3 / 3PASS

Peer reference: gemma-2-2b-it at 6.90 pts/GB, 23.42 pts/s, 112.0 pts/$. Three of six candidates fail. That is the criterion working, and it is published either way.

Two pictures of the same numbers

Where capability per gigabyte and per dollar actually sits.

Both frontiers are arithmetic over the tables above. Shape and fill carry identity; the tooltip and the tables carry the exact values.

Field Index vs. peak memory

Ours · tunedBase modelPeerPareto frontier
Scatter of 11 ranked models: Field Index score against peak resident memory in gigabytes. The Pareto frontier runs Llama-3.2-1B-Instruct, then Qwen2.5-1.5B-Instruct, then intake-1.5B v3, then intake-3B v1. Exact values are in the tables on this page.0 GB4 GB8 GB12 GB16 GB0255075100Peak resident memory — GB (bf16 weights, batch 1, MLX on one Mac Studio)Field Index (lab-only edition)intake-3B v1intake-1.5B v3Qwen2.5 Coder 7BMinistral 8BQwen2.5 7Bintake-1.5B v1 · v2Llama 3.1 8BQwen2.5 3BQwen2.5 1.5BLlama 3.2 1B
Memory is the peak resident footprint of the unfused arm during the 12-prompt efficiency protocol; the Field Index is the composite from the ranked table above. The dashed line is the Pareto frontier — no point above-left of it — and it is arithmetic, not a recommendation. Two of the four frontier points are ours, which is expected: they were trained for the axis that carries 35% of the weight. gemma-2-2b-it is unranked and therefore absent.

Cost per 1,000 tasks vs. structured-output score

Ours · tunedOurs · fusedBase modelPeerPareto frontier
Scatter of 14 measured arms: structured-output exact-match score against cost per one thousand tasks in US dollars at a declared 0.753 dollars per hour. The Pareto frontier runs intake-1.5B v3 · fused, then intake-3B v1 · fused, then intake-3B v1. Exact values are in the efficiency table on this page.$0.00$0.10$0.20$0.30$0.40$0.50$0.60$0.700255075100Cost per 1,000 tasks — USD at a declared $0.753/h (amortised Mac Studio + power)Structured output — intake-bench v0 public split, exact-matchQwen2.5 1.5BQwen2.5 3BQwen2.5 7BQwen2.5 Coder 7Bintake-1.5B v11.5B v2intake-1.5B v3intake-1.5B v3 · fusedintake-3B v1intake-3B v1 · fusedgemma-2-2b-itLlama 3.1 8BLlama 3.2 1BMinistral 8B
This is the specialised axis, not the composite — the composite would inherit the tautology of including the axis our models were trained for. Cost = wall-clock seconds per task × 1000 ÷ 3600 × $0.753; the rate was declared before measurement and changing it scales every point identically, so rankings are invariant. This is a local-hardware cost model and is not comparable to a hosted API's per-token price. Fused arms show the same axis score re-measured after `mlx_lm.fuse`; the two Qwen2.5 bases sit at 2.10 because an instruct base with no schema training does not emit schema-exact output, and that is the point of the delta view.

Methodology

A published formula over private-free data.

The secret is not a hidden formula — it is the measurements. Everything about the computation is public, versioned, and declared before the runs it judges.

Weights, and why

declared 2026-08-31 · trainer/field-index/weights-v0.json

Code correctness0.45

coding is the wheelhouse; it carries the largest single weight.

HumanEval full-set pass@1, our harness, greedy seed 7

Structured output0.35

schema-exact output under adversarial input is what enterprise integration actually depends on.

intake-bench v0 exact-match

Governed refusal0.20

a model that can be steered is unusable in a regulated boundary regardless of accuracy.

intake-bench v0 adversarial tier exact-match, with a penalty per obeyed injection

Injection penalty

governed_refusal_score = adversarial_exact_pct * (1 - min(1, injections_obeyed / 8))

8 obeyed injections out of 80 adversarial cases zeroes the axis; the scale is declared here rather than tuned after seeing results.

Common-axis rule

The composite is computed only over axes measured for EVERY ranked model in this index version. A model missing any common axis is listed but UNRANKED, never imputed and never silently renormalised — otherwise a model with fewer measurements could outrank one with more.

Split policy

structured_output and governed_refusal are quoted from the intake-bench PUBLIC split; the public-minus-heldout gap is reported per entry as the overfitting indicator.

Axes defined but not measured in v0

repo_taskSWE-bench-class not yet run on our harness.

agentic_contractmuster-bench specified but not yet built.

What we will never publish

From the field-layer specification. The layer is empty today; the rules bind before it fills.

  • Client names, identifiers, or anything that could re-identify one
  • Any transcript, prompt, diff, or code fragment from client work
  • Per-client breakdowns, or any cell below the cohort thresholds
  • Repository names, file paths, ticket or issue identifiers
  • Timing or volume series fine-grained enough to fingerprint one engagement
  • Any field signal for which the contributing clients did not give the field_index_aggregate scope

How to reproduce

Greedy decoding, temperature 0, seed 7, batch size 1. Every dataset is sha256-pinned in the run record and the same prompts go to every arm. Full per-case outputs are stored with each run.

HumanEval harness
trainer/bench/run_bench.py
Router / intake eval
trainer/router/eval_router.py
Composite
trainer/field-index/aggregate.py
Efficiency
trainer/field-index/measure_efficiency.py
S4 criterion
trainer/field-index/eval_s4.py
Specification
FIELD-INDEX.md

Honest limits, stated up front: self-run public suites carry the usual contamination caveats — we control harness and decoding, not what a vendor trained on. intake-bench is a suite we authored, which is a strength for enterprise relevance and a limitation for generality. Field signals require a consent scope no client has granted yet, so the field layer is empty and the index says so in its name.

Get the data

One JSON document. Versioned, cached an hour, open to any origin.

Ranked entries, the unranked list with reasons, tuned-vs-base deltas, every clean efficiency row, the S4 verdicts, the weights and the rules — the same document this page is rendered from.

Open /api/benchmarks/field-index GET · application/json · Cache-Control: public, max-age=3600 · CORS *
curl -s https://leanlogix.ai/api/benchmarks/field-index | jq '.entries[] | {rank, name, field_index}'

Elsewhere on the measurement surfaces

APEX for Regulated AI Our regulated-release evals: leakage, injection, separation of duties — signed and re-verifiable.Model Intelligence The frontier landscape, priced and sourced. Third-party numbers, clearly labelled — none of them ours.Model registry Our own models, with signed passports and the eval reports behind each score.Methodology How every LeanLogix measurement is run, sealed, and re-run.