Skip to content
Benchmark evidence · reviewed 2026-09-25

Model benchmarks and costs

The default view compares GPT-6 Sol and Luna at API max on the September 25 Intelligence Index v4.3.2 snapshot. Choose another benchmark to compare models within that version. Each score links to its source; costs show whether they come from the published task or a stated token estimate. API max results do not establish Codex CLI max support.

See our own controlled studies for local runs over versioned repository fixtures; this matrix instead compares external published scores joined to dated prices.

Opus 5, Fable 5.1 and Astra: what the measurements support

Start with published benchmarks, then use AGT and controlled local trials to check transfer to your work. These assessments retain each benchmark’s score, effort and test conditions.

Loading model assessments…

Model × effort matrix

Compare within one benchmark, version and harness. Scores are point estimates; an overlapping interval does not establish a reliable quality difference. Snapshot dates mark when a public row was observed, not when its trials ran. AGT observations are an additional local signal.

Loading the evidence-backed answer…

Loading generated matrix data…

At least the reference score, lower price

A candidate must meet the reference point estimate on the selected score and cheaper on published task cost or on the declared token-price mix. This library recommends candidates; it does not route work or lower correctness floors.

Snapshot provenance