Which models exist, what they cost, where each one leads, how to read the boards without being misled, and how a shortlist becomes a governed release. Every number below is a published, third-party result with its source attached — none of them are LeanLogix measurements, and we say which is which.
As of August 2026·9 models charted · 14 boards snapshotted·every figure sourced · aggregator disagreements flagged
Cost vs. intelligence
The landscape in one view.
Intelligence-index score against blended list price. The dashed line is the price-performance frontier — no model above the line is both cheaper and stronger than a model on it. The most expensive model is not the strongest, and the frontier includes open weights you can run inside your own boundary.
Index scores are tracker snapshots of the Artificial Analysis Intelligence Index (v4.1.x) and shift with effort settings; blended price is our stated 3:1 convention over list prices, not a billed price. Models missing either sourced number are excluded rather than estimated.
Who leads what — and what each board actually measures.
Representative leaders per board, not an exhaustive matrix — the linked live boards are authoritative. Where aggregators disagree, the card says so instead of picking a winner.
Aggregate indices
Artificial Analysis Intelligence Index
Weighted composite of 9 evals across agents, coding, general capability, and science.
Claude Opus 5Anthropic
63.0
Claude Fable 5Anthropic
62.1
Kimi K3Moonshot AIOPENtop open-weights
59.7
GPT-5.6 SolOpenAI
58.9
Qwen3.8 Max PreviewAlibaba
58.1
Claude Opus 4.8Anthropic
57.3
0–100 weighted composite, re-weighted toward agentic work in v4.1 (10 evals — GDPval-AA, Terminal-Bench, GPQA Diamond, SciCode and others). One snapshot, one effort setting: AA's own v4.1.1 article lists Opus 5 at 60.7 and Fable 5 at 59.9 at a lower effort setting, so compare models within a setting, never across.
BenchLM snapshot of the AA Intelligence Index (11 Aug 2026)live board
Human preference
LMArena (Chatbot Arena)
Blind A/B human votes across 360+ models — Elo from 6.8M+ votes.
Claude Fable 5Anthropic
1506 ±5 Elo
Claude Opus 4.6 ThinkingAnthropic
1505 ±4 Elo
Claude Opus 4.7 ThinkingAnthropic
1502 ±4 Elo
Muse Spark 1.2 (xHigh)Meta
1499 ±10 Elo
Claude Opus 4.6Anthropic
1498 ±3 Elo
Read straight off the live board: the top five sit inside 8 Elo, narrower than their own confidence intervals — one tier, not a ranking. Secondary trackers still quote ~1525 for the leader against the pre-re-baseline scale. Newest releases take weeks of votes to land, so this board trails the aggregate indices.
Arena (arena.ai) text leaderboard, read 11 Aug 2026live board
Knowledge & reasoning
GPQA Diamond
Graduate-level science questions experts get wrong ~60% of the time even with web access.
Sakana Fugu-UltraSakana AI
95.5%
Sakana FuguSakana AI
95.5%
GPT-5.6 SolOpenAI94.1% on EdenAI
94.6%
Gemini 3.1 ProGoogle
94.3%
Claude Opus 4.7 (Adaptive)Anthropic
94.2%
Effectively saturated: the whole top ten spans 2.7 points, so the ordering is inside the noise. Scoring policy now moves the number more than capability does — one 2026 dispute had the same model at 93.2% or 55.6% depending on whether safety refusals count as correct. Always check the protocol.
BenchLM (11 Aug 2026) — sources disagreelive board
Code & software engineering
SWE-bench Verified
Resolving real GitHub issues in real repositories (human-verified subset).
Claude Opus 5Anthropic97.0% on llm-stats — trackers disagree
Saturating: the top four sit within ~1 point and trackers disagree by a full point on the leader. Anthropic holds six of the top six on the BenchLM snapshot; the best open-weights entry is 82.4%, with DeepSeek V4 Pro (Max) just behind at 80.6%. The separating signal has moved to SWE-bench Pro.
Novel abstract-reasoning puzzles designed to resist memorization.
GPT-5.6 SolOpenAI
92.5%
Claude Opus 5Anthropic
90.4%
GPT-5.5OpenAI
85%
Nearing saturation. The frontier has moved to ARC-AGI-3, where Claude Opus 5 leads at 30.2% — roughly 3× the next model — so the fluid-reasoning race is wide open again.
Thousands of expert-authored questions at the frontier of human knowledge.
Claude Opus 5Anthropic
64.7%
Claude Mythos 5Anthropic
64.5%
Muse Spark 1.1Meta
62.1%
GPT-5.4 ProOpenAI
58.7%
Claude Opus 4.8Anthropic
57.9%
Tool-assisted and closed-book HLE runs are effectively answering different questions, and tool-assisted lands 10+ points higher — pricepertoken puts Claude Fable 5 at 55.5% closed-book against these tool-assisted figures. Treat small gaps as directional, not decisive, and always check the protocol.
BenchLM (11 Aug 2026) — protocol varies by sourcelive board
Code & software engineering
LiveCodeBench
Fresh competitive-programming problems collected after model cutoffs (contamination-resistant).
Tool-Agent-User interaction in realistic domains (retail, airline, banking).
Step-3.5-FlashStepFun
0.882
GPT-5.6 SolOpenAIτ²-Bench, Artificial Analysis
85.1%
Small and mid-size models routinely punch above their weight on tool-agent-user interaction — a reminder that agentic fit is task-shaped, not size-shaped.
Research-grade math problems authored by professional mathematicians.
GPT-5.6 SolOpenAI
89.0%
GPT-5.6 TerraOpenAI
84.9%
GPT-5.6 LunaOpenAI
78.6%
GPT-5.5 ProOpenAI
52.4%
FrontierMath (legacy split), kept for historical reference. OpenAI sweeps research-grade math, and the 78.6% → 52.4% drop between generations is the steepest single-generation gap on any board here.
Olympiad-level competition math (American Invitational Mathematics Exam).
GLM-5.2Z AI / ZhipuOPEN
99.2%
InklingThinking Machines Lab
97.1%
Kimi K2.6Moonshot AIOPEN
96.4%
GLM-5Z AI / ZhipuOPEN
95.8%
Competition math is saturated — the top models sit within 2.8 points, and open-weights hold the top slot outright. The signal has moved to FrontierMath.
Unchanged this week, and stale by design: the official board has not been re-run since 20 Nov 2025, so no 2026 frontier release appears in it at all. 225 Exercism exercises across six languages. llm-stats' mirror reports a different DeepSeek V3.2-Exp figure (0.745) for a different variant — read this board as a 2025 baseline, not a current ranking.
Aider official leaderboard (last re-run 20 Nov 2025)live board
Code & software engineeringSATURATED
HumanEval
Functional correctness of generated Python (pass@1 against unit tests).
A historical LeanLogix endpoint measurement on a small, saturated sample. It is evidence, not a current availability, boundary, or loaded-weight attestation.
LeanLogix (measured) + published frontierlive board
Reading the boards
How to read a leaderboard in 2026 without being misled.
Contamination is the default assumption
Static suites (MMLU, GSM8K, HumanEval) are saturated or leaked into training data. Prefer contamination-guarded boards — LiveCodeBench collects problems after model cutoffs; LiveBench replaces ~1/6 of questions monthly.
Harness choice alone swings scores 10–20 points; HLE diverges 10+ points between tool-assisted and closed-book runs; the same Terminal-Bench name shows different leaders on different harnesses. Compare within one harness only.
Arena's top tier sits inside ~55 Elo, and top SWE-bench Verified scores sit within ~1 point — inside the confidence interval, ordering is noise. Treat the cluster as one tier.
The Leaderboard Illusion analysis (2M battles) documented undisclosed private variant testing and preferential sampling on arena-style boards; pairwise voting also rewards confident-wrong over hedged-correct answers.
Practitioner consensus: leaderboard position has little correlation with whether a model works on your task. The leaderboard is the shortlist; your own evaluation is the decision.
The list prices, then the levers that decide what you actually pay. Sticker $/Mtok routinely misses the real bill by an order of magnitude once output ratios, caching, and agent loops are in play.
Models differ several-fold in output tokens spent per task solved — one Aug 2026 comparison measured a 4× spread on the same agentic benchmark — so $/task, not $/token, is the number to compare.
Several majors re-price above a context threshold (e.g. higher rates past ~200–272K input tokens) — long-document workloads should be priced at the surcharge tier.
Gateway usage data shows production traffic concentrating in 'good enough and far cheaper' models rather than leaderboard leaders — benchmark rank and deployed share are nearly uncorrelated.
The leaderboard is the shortlist. Your evaluation is the decision.
Five steps from 300+ models to one you can defend — the last two are where LeanLogix does its real work.
01
Shortlist from the aggregate boards
Use the intelligence indexes and category boards to cut 300+ models to a handful — reading clusters as tiers, at one effort setting, within one harness.
Model the real bill: input:output ratio, cache hit rate, agent-loop depth, context tier, tokens-per-task efficiency. A cheaper $/Mtok model can cost more per task solved.
Regulated or private data pushes toward open weights inside your boundary; latency and unit economics push toward small models; context needs and modality trim the list again.
A few hundred examples from your real task, scored by deterministic checkers — not an LLM judge. This is the decision; everything above was the shortlist.
Bind the chosen artifact to local evidence, independent approval, and a signed release record — the part regulators now ask for, with EU GPAI enforcement live since 2 Aug 2026.
Every way a model gets better — and when each one earns its cost.
2026 is the year of tuned small models: adapter methods over open bases now cover most enterprise specialization, with distillation carrying frontier behavior down-market. What follows is the toolbox — what each method is, when it applies, and what it costs. Our own governed training surface documents which of these LeanLogix has evidence for.
PretrainData-center scale · trillions of tokens
Pretraining from scratch
Learning a base model from raw corpora — the frontier-lab lane, now at ~10²⁵–10²⁶ FLOP per run.
When: Almost never the enterprise answer: the EU AI Act attaches systemic-risk duties above 10²⁵ FLOP, and open bases already cover most starting points.
Fine-tuneHighest VRAM · full checkpoint storage
Full fine-tune
Updating every weight on your task data.
When: Rarely worth it for narrow tasks — our own cycle evidence found full fine-tuning no better than LoRA on an 18-probe suite, and reverted it.
Fine-tuneSingle-GPU / Apple-Silicon class · adapter is MBs, not GBs
LoRA / QLoRA
Low-rank adapters beside frozen weights; QLoRA trains them over a 4-bit base.
When: The 2026 default for in-boundary specialization — small artifact, auditable diff, trains on workstation-class hardware.
Fine-tune≈ LoRA memory · slightly longer runs
DoRA / GaLore
Weight-decomposed adapters and gradient low-rank projection — closer-to-full-rank quality at adapter-like memory.
When: When LoRA plateaus below the quality bar and you still can't afford full-rank training.
AlignComparable to fine-tuning · needs preference pairs
DPO / GRPO
Preference optimization and grouped-rollout RL directly on comparison data — no separate reward model.
When: After supervised tuning, to sharpen refusal behavior, formatting discipline, and reasoning-under-policy.
AlignCompute-heavy · labeler-light
RLAIF / self-rewarding
AI-generated feedback replacing human labels in the alignment loop.
When: To scale alignment data past what labelers can produce — with the judge itself now a component you must evaluate.
A large teacher grades or generates targets for a small student on the student's own outputs.
When: The frontier-to-small pipeline: when a tuned small model must approach big-model behavior on one domain at a fraction of the serving cost.
CompressOne-shot conversion · ~4× smaller at 4-bit
Quantization (GGUF / AWQ / GPTQ / FP8)
Lower-precision weights for deployment — 4-bit typically retains ~97–99% quality (published ranges, format-dependent).
When: At deploy time, to fit the boundary's hardware; retention must be measured on your evals, not assumed from the format's reputation.
CompressOne-shot · quality re-check required
Pruning (Wanda / SparseGPT)
One-shot removal of low-salience weights.
When: When quantization alone doesn't reach the footprint target and you can re-verify quality after.
ComposeMinutes of compute · no training data needed
Merging (TIES / DARE / SLERP)
Combining adapters or variants in weight space without retraining.
When: To compose capabilities from separately trained adapters — cheap, but provenance of the merged artifact must stay explicit.
LoopThe discipline, not the compute
Continual eval-gated loop
Measure → modify → keep or discard, with every kept change justified by an eval delta.
When: Always — it is the difference between a training run and a training process. Each cycle's verdict (kept / reverted / inconclusive) is release evidence.
For what LeanLogix itself has training evidence for — and what is deliberately not offered as a runnable lane — see Training evidence & roadmap.
What this page is — and is not
Everything above is published, third-party data with its source attached — tracker snapshots and list prices as of August 2026, refreshed on a stated date rather than continuously. Aggregators disagree; where they do, the disagreement is shown instead of resolved. None of these numbers are LeanLogix measurements: our own models are scored on our own hard suite at /leaderboard, and the two are never mixed. In the portfolio, upstream artifact selection and deployment-fit authority remain with MOEModels; this page is a sourced reading of the public landscape, not a procurement recommendation or a release decision.