Learn

Field notes from training models that have to be defended.

We train small models on one Mac Studio, hold them to criteria we write down before the run, and publish the verdicts either way. These guides are built from those runs: the adapter that forgot Python, the efficiency claim that came back a quarter of its size, the router that answers coding problems with nothing at all. They are for engineers who have to pick a model, measure it, tune it, or combine it, and then explain the number to someone who can check.

  1. Benchmarking6 min read

    HumanEval 57.93 to 3.05. The gate refused, and we kept the receipt.

    Our own publication gate turned down our own model, the decision memo told the founder no, and a rename broke a signature chain. Why a lab that prints its failing numbers is the one whose passing numbers are worth reading.

    Read the guide
  2. Choosing6 min read

    We declared an order of magnitude. The measurement came back 2.16×.

    The efficiency claim was written down before anyone ran it, and the number that came back embarrassed the claim. What went wrong was the choice of opponent, and it is the same mistake most small-model comparisons make.

    Read the guide
  3. Training7 min read

    Cycle three, and the failure moved again.

    Three training cycles on a 1.5B intake specialist, three different ways of losing general ability, and a 3B run that refuted the obvious fix. What LoRA, distillation, and preference tuning actually cost at this size, and why the stopping rule is written before the run.

    Read the guide
  4. Benchmarking6 min read

    Two peers tied at 34.03, and the verdict flipped on which one we picked.

    A one-line ambiguity in a criterion we wrote ourselves would have let us choose the flattering comparison. What has to be decided before a benchmark number exists, and why the open-model seat has been empty since March 2025.

    Read the guide
  5. Composing6 min read

    The router that answered 164 coding problems with silence.

    Bearing routes at 96.00 and emits nothing on HumanEval. That is the right behaviour in front of a specialist and the wrong behaviour as one. Routers, weight merges, and upcycled mixtures each answer the specialist's problem differently, and each owes different evidence.

    Read the guide

The measurements behind these guides are on the Field Index, and the method is on the methodology page. Longer arguments about regulated AI are in Insights.