Evidence · Frontier landscape

The model landscape, priced and sourced.

Which models exist, what they cost, where each one leads, how to read the boards without being misled, and how a shortlist becomes a governed release. Every number below is a published, third-party result with its source attached — none of them are LeanLogix measurements, and we say which is which.

As of August 20269 models charted · 14 boards snapshottedevery figure sourced · aggregator disagreements flagged

Cost vs. intelligence

The landscape in one view.

Intelligence-index score against blended list price. The dashed line is the price-performance frontier — no model above the line is both cheaper and stronger than a model on it. The most expensive model is not the strongest, and the frontier includes open weights you can run inside your own boundary.

Closed weightsOpen weightsPrice-performance frontier
$2$3$5$10$2045505560Blended list price — USD per 1M tokens (3:1 in:out, log scale)Artificial Analysis Intelligence IndexClaude Opus 5 (Anthropic) — index 60.7, $5/$25 per Mtok (blended $10.00)Claude Opus 5Claude Fable 5 (Anthropic) — index 59.9, $10/$50 per Mtok (blended $20.00)Claude Fable 5GPT-5.6 Sol (OpenAI) — index 58.9, $5/$30 per Mtok (blended $11.25)GPT-5.6 SolKimi K3 (Moonshot AI) — index 57, $3/$15 per Mtok (blended $6.00)Kimi K3Grok 4.5 (xAI) — index 54, $2/$6 per Mtok (blended $3.00)Grok 4.5GLM-5.2 (Z AI / Zhipu) — index 51, $1.4/$4.4 per Mtok (blended $2.15)GLM-5.2Muse Spark 1.1 (Meta) — index 51, $1.25/$4.25 per Mtok (blended $2.00)Muse Spark 1.1Gemini 3.6 Flash (Google) — index 50, $1.5/$7.5 per Mtok (blended $3.00)Gemini 3.6 FlashGemini 3.1 Pro (Google) — index 48, $2/$12 per Mtok (blended $4.50)Gemini 3.1 Pro
Index scores are tracker snapshots of the Artificial Analysis Intelligence Index (v4.1.x) and shift with effort settings; blended price is our stated 3:1 convention over list prices, not a billed price. Models missing either sourced number are excluded rather than estimated.
ModelOrgIndex$ in / Mtok$ out / MtokBlendedWeightsSource
Muse Spark 1.1FRONTIERMeta51$1.25$4.25$2.00closedsource
GLM-5.2Z AI / Zhipu51$1.4$4.4$2.15opensource
Grok 4.5FRONTIERxAI54$2$6$3.00closedsource
Gemini 3.6 FlashGoogle50$1.5$7.5$3.00closedsource
Gemini 3.1 ProGoogle48$2$12$4.50closedsource
Kimi K3FRONTIERMoonshot AI57$3$15$6.00opensource
Claude Opus 5FRONTIERAnthropic60.7$5$25$10.00closedsource
GPT-5.6 SolOpenAI58.9$5$30$11.25closedsource
Claude Fable 5Anthropic59.9$10$50$20.00closedsource

Benchmark boards · August 2026

Who leads what — and what each board actually measures.

Representative leaders per board, not an exhaustive matrix — the linked live boards are authoritative. Where aggregators disagree, the card says so instead of picking a winner.

Aggregate indices

Artificial Analysis Intelligence Index

Weighted composite of 9 evals across agents, coding, general capability, and science.

Claude Opus 5Anthropic
63.0
Claude Fable 5Anthropic
62.1
Kimi K3Moonshot AIOPENtop open-weights
59.7
GPT-5.6 SolOpenAI
58.9
Qwen3.8 Max PreviewAlibaba
58.1
Claude Opus 4.8Anthropic
57.3

0–100 weighted composite, re-weighted toward agentic work in v4.1 (10 evals — GDPval-AA, Terminal-Bench, GPQA Diamond, SciCode and others). One snapshot, one effort setting: AA's own v4.1.1 article lists Opus 5 at 60.7 and Fable 5 at 59.9 at a lower effort setting, so compare models within a setting, never across.

BenchLM snapshot of the AA Intelligence Index (11 Aug 2026)live board
Human preference

LMArena (Chatbot Arena)

Blind A/B human votes across 360+ models — Elo from 6.8M+ votes.

Claude Fable 5Anthropic
1506 ±5 Elo
Claude Opus 4.6 ThinkingAnthropic
1505 ±4 Elo
Claude Opus 4.7 ThinkingAnthropic
1502 ±4 Elo
Muse Spark 1.2 (xHigh)Meta
1499 ±10 Elo
Claude Opus 4.6Anthropic
1498 ±3 Elo

Read straight off the live board: the top five sit inside 8 Elo, narrower than their own confidence intervals — one tier, not a ranking. Secondary trackers still quote ~1525 for the leader against the pre-re-baseline scale. Newest releases take weeks of votes to land, so this board trails the aggregate indices.

Arena (arena.ai) text leaderboard, read 11 Aug 2026live board
Knowledge & reasoning

GPQA Diamond

Graduate-level science questions experts get wrong ~60% of the time even with web access.

Sakana Fugu-UltraSakana AI
95.5%
Sakana FuguSakana AI
95.5%
GPT-5.6 SolOpenAI94.1% on EdenAI
94.6%
Gemini 3.1 ProGoogle
94.3%
Claude Opus 4.7 (Adaptive)Anthropic
94.2%

Effectively saturated: the whole top ten spans 2.7 points, so the ordering is inside the noise. Scoring policy now moves the number more than capability does — one 2026 dispute had the same model at 93.2% or 55.6% depending on whether safety refusals count as correct. Always check the protocol.

BenchLM (11 Aug 2026) — sources disagreelive board
Code & software engineering

SWE-bench Verified

Resolving real GitHub issues in real repositories (human-verified subset).

Claude Opus 5Anthropic97.0% on llm-stats — trackers disagree
96.0%
GPT-5.6 SolOpenAI
96.2%
Claude Mythos 5Anthropic
95.5%
Claude Fable 5Anthropic
95.0%
Ornith-1.0-397BDeepReinforce AIOPENtop open-weights (#8)
82.4%

Saturating: the top four sit within ~1 point and trackers disagree by a full point on the leader. Anthropic holds six of the top six on the BenchLM snapshot; the best open-weights entry is 82.4%, with DeepSeek V4 Pro (Max) just behind at 80.6%. The separating signal has moved to SWE-bench Pro.

BenchLM (11 Aug 2026) / llm-statslive board
Code & software engineering

SWE-bench Pro

Harder, contamination-guarded successor to SWE-bench — longer-horizon, multi-file repository work.

Claude Fable 5Anthropic
80.3%
Grok 4.5xAI
64.7%
GPT-5.6 SolOpenAI
64.6%
Claude Sonnet 5Anthropic
63.2%
GLM-5.2Z AI / ZhipuOPENtop open-weights
62.1%

Where Verified compresses to ~1 point, Pro still separates the frontier by 15+ — the current signal for long-horizon software engineering.

Tracker roundups (Aug 2026)live board
Knowledge & reasoning

ARC-AGI-2

Novel abstract-reasoning puzzles designed to resist memorization.

GPT-5.6 SolOpenAI
92.5%
Claude Opus 5Anthropic
90.4%
GPT-5.5OpenAI
85%

Nearing saturation. The frontier has moved to ARC-AGI-3, where Claude Opus 5 leads at 30.2% — roughly 3× the next model — so the fluid-reasoning race is wide open again.

BenchLM / ARC Prize (Aug 2026)live board
Agentic & tool use

Terminal-Bench

Completing real tasks in a live terminal/sandbox.

GPT-5.6 SolOpenAITerminal-Bench 2.0
91.9%
Claude Mythos 5Anthropic
88.0%
GPT-5.6 TerraOpenAI
87.4%

Strongly harness-dependent: llm-stats' own Terminal-Bench 2 harness shows a different leader across 49 models. Compare within one harness only.

BenchLM (Aug 2026) — harness-dependentlive board
Knowledge & reasoning

Humanity's Last Exam (HLE)

Thousands of expert-authored questions at the frontier of human knowledge.

Claude Opus 5Anthropic
64.7%
Claude Mythos 5Anthropic
64.5%
Muse Spark 1.1Meta
62.1%
GPT-5.4 ProOpenAI
58.7%
Claude Opus 4.8Anthropic
57.9%

Tool-assisted and closed-book HLE runs are effectively answering different questions, and tool-assisted lands 10+ points higher — pricepertoken puts Claude Fable 5 at 55.5% closed-book against these tool-assisted figures. Treat small gaps as directional, not decisive, and always check the protocol.

BenchLM (11 Aug 2026) — protocol varies by sourcelive board
Code & software engineering

LiveCodeBench

Fresh competitive-programming problems collected after model cutoffs (contamination-resistant).

Gemini 3 Pro PreviewGoogle
91.7%
Gemini 3 Flash PreviewGoogle
90.8%
DeepSeek V3.2 SpecialeDeepSeekOPENtop open-weights
89.6%

Contamination-resistant by design (post-cutoff problems), which is exactly why its ordering differs from the static coding boards.

pricepertoken (1 Aug 2026 snapshot)live board
Agentic & tool use

τ-bench (tau-bench)

Tool-Agent-User interaction in realistic domains (retail, airline, banking).

Step-3.5-FlashStepFun
0.882
GPT-5.6 SolOpenAIτ²-Bench, Artificial Analysis
85.1%

Small and mid-size models routinely punch above their weight on tool-agent-user interaction — a reminder that agentic fit is task-shaped, not size-shaped.

llm-stats (Aug 2026)live board
Math

FrontierMath

Research-grade math problems authored by professional mathematicians.

GPT-5.6 SolOpenAI
89.0%
GPT-5.6 TerraOpenAI
84.9%
GPT-5.6 LunaOpenAI
78.6%
GPT-5.5 ProOpenAI
52.4%

FrontierMath (legacy split), kept for historical reference. OpenAI sweeps research-grade math, and the 78.6% → 52.4% drop between generations is the steepest single-generation gap on any board here.

BenchLM (11 Aug 2026)live board
Math

AIME 2026

Olympiad-level competition math (American Invitational Mathematics Exam).

GLM-5.2Z AI / ZhipuOPEN
99.2%
InklingThinking Machines Lab
97.1%
Kimi K2.6Moonshot AIOPEN
96.4%
GLM-5Z AI / ZhipuOPEN
95.8%

Competition math is saturated — the top models sit within 2.8 points, and open-weights hold the top slot outright. The signal has moved to FrontierMath.

BenchLM (11 Aug 2026)live board
Code & software engineering

Aider Polyglot

Editing existing code correctly across many languages via diffs.

GPT-5 (high)OpenAI
88.0%
GPT-5 (medium)OpenAI
86.7%
o3-pro (high)OpenAI
84.9%
DeepSeek-V3.2-Exp (Chat)DeepSeekOPENtop open-weights
70.2%

Unchanged this week, and stale by design: the official board has not been re-run since 20 Nov 2025, so no 2026 frontier release appears in it at all. 225 Exercism exercises across six languages. llm-stats' mirror reports a different DeepSeek V3.2-Exp figure (0.745) for a different variant — read this board as a 2025 baseline, not a current ranking.

Aider official leaderboard (last re-run 20 Nov 2025)live board
Code & software engineeringSATURATED

HumanEval

Functional correctness of generated Python (pass@1 against unit tests).

SprintLoop-7B · v6LeanLogixOPENOURS · MEASURED7B, in-boundary — measured on our endpoint, n=30
93.3% pass@1
Frontier (GPT-5.5 / Opus 4.8)OpenAI / Anthropic
~99%
Qwen2.5-Coder-7B (our base)AlibabaOPEN
~88%

A historical LeanLogix endpoint measurement on a small, saturated sample. It is evidence, not a current availability, boundary, or loaded-weight attestation.

LeanLogix (measured) + published frontierlive board

Reading the boards

How to read a leaderboard in 2026 without being misled.

Contamination is the default assumption

Static suites (MMLU, GSM8K, HumanEval) are saturated or leaked into training data. Prefer contamination-guarded boards — LiveCodeBench collects problems after model cutoffs; LiveBench replaces ~1/6 of questions monthly.

The harness moves the number

Harness choice alone swings scores 10–20 points; HLE diverges 10+ points between tool-assisted and closed-book runs; the same Terminal-Bench name shows different leaders on different harnesses. Compare within one harness only.

Read clusters, not ranks

Arena's top tier sits inside ~55 Elo, and top SWE-bench Verified scores sit within ~1 point — inside the confidence interval, ordering is noise. Treat the cluster as one tier.

Preference boards have structural bias

The Leaderboard Illusion analysis (2M battles) documented undisclosed private variant testing and preferential sampling on arena-style boards; pairwise voting also rewards confident-wrong over hedged-correct answers.

Rank does not predict production fit

Practitioner consensus: leaderboard position has little correlation with whether a model works on your task. The leaderboard is the shortlist; your own evaluation is the decision.

Cost, performance, size

Price the workload, not the token.

The list prices, then the levers that decide what you actually pay. Sticker $/Mtok routinely misses the real bill by an order of magnitude once output ratios, caching, and agent loops are in play.

ModelOrg$ in / Mtok$ out / MtokContextWeightsNotesSource
Claude Opus 5Anthropic$5$251Mclosedoptional fast mode at $10/$50source
Claude Fable 5Anthropic$10$50~1M (third-party reported)closedsource
Claude Sonnet 5Anthropic$2$101Mclosedintro pricing through 31 Aug 2026, then $3/$15source
Claude Haiku 4.5Anthropic$1$5200Kclosedsource
GPT-5.6 SolOpenAI$5$301.1Mclosed$10/$45 above 272K input tokenssource
GPT-5.6 TerraOpenAI$2$121.1Mclosedafter the 30 Jul 2026 price cut (−20%)source
GPT-5.6 LunaOpenAI$0.2$1.21.1Mclosedafter the 30 Jul 2026 price cut (−80%)source
Gemini 3.1 ProGoogle$2$121Mclosed$4/$18 above 200K context; cached input $0.20source
Gemini 3.6 FlashGoogle$1.5$7.51Mclosedcached input $0.15source
Grok 4.5xAI$2$6500Kclosednotably token-efficient on agentic taskssource
Qwen3.8-MaxAlibaba$2$61Mclosedopen weights announced as forthcoming, not yet publishedsource
Kimi K3Moonshot AI$3$151MOPENlargest open model published (2.8T total / 104B active)source
GLM-5.2Z AI / Zhipu$1.4$4.41MOPENMIT license, 753B MoE / 40B activesource
DeepSeek V4 ProDeepSeek$0.435$0.871MclosedV4 weights release not confirmed in tracked sourcessource
Mistral Medium 3.5Mistral$1.5$7.5128KOPENdense 128B, Modified MIT — self-hostable on 4 GPUssource
MiniMax M3MiniMax$0.3$1.21MOPENsource

Output ≠ input

Output tokens list at 4–8× input across the majors, so a generation-heavy workload prices completely differently from a summarization one.

Prompt caching

Cached repeated input is discounted up to ~90% — a long, stable system prompt changes the math entirely once caching is on.

Agent loops multiply everything

A single coding-agent task consumes 400K–2M cumulative input tokens across its loop; retries and tool-call fan-out multiply requests again.

Token efficiency varies by model

Models differ several-fold in output tokens spent per task solved — one Aug 2026 comparison measured a 4× spread on the same agentic benchmark — so $/task, not $/token, is the number to compare.

Long-context surcharges

Several majors re-price above a context threshold (e.g. higher rates past ~200–272K input tokens) — long-document workloads should be priced at the surcharge tier.

Deployed demand follows price-performance

Gateway usage data shows production traffic concentrating in 'good enough and far cheaper' models rather than leaderboard leaders — benchmark rank and deployed share are nearly uncorrelated.

From landscape to release

The leaderboard is the shortlist. Your evaluation is the decision.

Five steps from 300+ models to one you can defend — the last two are where LeanLogix does its real work.

01

Shortlist from the aggregate boards

Use the intelligence indexes and category boards to cut 300+ models to a handful — reading clusters as tiers, at one effort setting, within one harness.

This page
02

Price your workload shape, not the token

Model the real bill: input:output ratio, cache hit rate, agent-loop depth, context tier, tokens-per-task efficiency. A cheaper $/Mtok model can cost more per task solved.

The price book
03

Size for the boundary

Regulated or private data pushes toward open weights inside your boundary; latency and unit economics push toward small models; context needs and modality trim the list again.

Footprint reference
04

Run your own deterministic evals

A few hundred examples from your real task, scored by deterministic checkers — not an LLM judge. This is the decision; everything above was the shortlist.

Our methodology
05

Govern the release

Bind the chosen artifact to local evidence, independent approval, and a signed release record — the part regulators now ask for, with EU GPAI enforcement live since 2 Aug 2026.

Release evidence

The training toolbox

Every way a model gets better — and when each one earns its cost.

2026 is the year of tuned small models: adapter methods over open bases now cover most enterprise specialization, with distillation carrying frontier behavior down-market. What follows is the toolbox — what each method is, when it applies, and what it costs. Our own governed training surface documents which of these LeanLogix has evidence for.

PretrainData-center scale · trillions of tokens

Pretraining from scratch

Learning a base model from raw corpora — the frontier-lab lane, now at ~10²⁵–10²⁶ FLOP per run.

When: Almost never the enterprise answer: the EU AI Act attaches systemic-risk duties above 10²⁵ FLOP, and open bases already cover most starting points.

Fine-tuneHighest VRAM · full checkpoint storage

Full fine-tune

Updating every weight on your task data.

When: Rarely worth it for narrow tasks — our own cycle evidence found full fine-tuning no better than LoRA on an 18-probe suite, and reverted it.

Fine-tuneSingle-GPU / Apple-Silicon class · adapter is MBs, not GBs

LoRA / QLoRA

Low-rank adapters beside frozen weights; QLoRA trains them over a 4-bit base.

When: The 2026 default for in-boundary specialization — small artifact, auditable diff, trains on workstation-class hardware.

Fine-tune≈ LoRA memory · slightly longer runs

DoRA / GaLore

Weight-decomposed adapters and gradient low-rank projection — closer-to-full-rank quality at adapter-like memory.

When: When LoRA plateaus below the quality bar and you still can't afford full-rank training.

AlignComparable to fine-tuning · needs preference pairs

DPO / GRPO

Preference optimization and grouped-rollout RL directly on comparison data — no separate reward model.

When: After supervised tuning, to sharpen refusal behavior, formatting discipline, and reasoning-under-policy.

AlignCompute-heavy · labeler-light

RLAIF / self-rewarding

AI-generated feedback replacing human labels in the alignment loop.

When: To scale alignment data past what labelers can produce — with the judge itself now a component you must evaluate.

DistillTeacher inference dominates · student trains cheap

On-policy distillation

A large teacher grades or generates targets for a small student on the student's own outputs.

When: The frontier-to-small pipeline: when a tuned small model must approach big-model behavior on one domain at a fraction of the serving cost.

CompressOne-shot conversion · ~4× smaller at 4-bit

Quantization (GGUF / AWQ / GPTQ / FP8)

Lower-precision weights for deployment — 4-bit typically retains ~97–99% quality (published ranges, format-dependent).

When: At deploy time, to fit the boundary's hardware; retention must be measured on your evals, not assumed from the format's reputation.

CompressOne-shot · quality re-check required

Pruning (Wanda / SparseGPT)

One-shot removal of low-salience weights.

When: When quantization alone doesn't reach the footprint target and you can re-verify quality after.

ComposeMinutes of compute · no training data needed

Merging (TIES / DARE / SLERP)

Combining adapters or variants in weight space without retraining.

When: To compose capabilities from separately trained adapters — cheap, but provenance of the merged artifact must stay explicit.

LoopThe discipline, not the compute

Continual eval-gated loop

Measure → modify → keep or discard, with every kept change justified by an eval delta.

When: Always — it is the difference between a training run and a training process. Each cycle's verdict (kept / reverted / inconclusive) is release evidence.

For what LeanLogix itself has training evidence for — and what is deliberately not offered as a runnable lane — see Training evidence & roadmap.

What this page is — and is not

Everything above is published, third-party data with its source attached — tracker snapshots and list prices as of August 2026, refreshed on a stated date rather than continuously. Aggregators disagree; where they do, the disagreement is shown instead of resolved. None of these numbers are LeanLogix measurements: our own models are scored on our own hard suite at /leaderboard, and the two are never mixed. In the portfolio, upstream artifact selection and deployment-fit authority remain with MOEModels; this page is a sourced reading of the public landscape, not a procurement recommendation or a release decision.

Live boards & sourcesArtificial AnalysisArena (LMArena)OpenRouter — pricing & usagellm-statsBenchLMARC PrizeLiveBenchEU AI Act implementation timeline
How we evaluate for release Our measured scores Seal a route decision