Skip to content
Policy revision · 25 September 2026

The core Luna and Sol routes move to GPT-6

New fleet cards prefer GPT-6 Sol. This operator decision is backed by dated prices and vendor claims; local completion evidence remains provisional.

Task-class candidates and fallbacks · External benchmark matrix · Authoritative policy and cohort plan

Before and after

Standard USD per million input / cached-input / output tokens, dated 25 September 2026. Equal token prices do not establish equal cost per completed card.

Loading routing policy…
Retain Terra temporarily. GPT-5.6 Terra keeps the 21–50 score band and the one-tier reissue step. Its input/cache prices match GPT-6 Sol; output costs 20% more. Astra costs five times Sol and has unresolved access reliability. No score threshold or correctness floor changes.

Price history and sources. GPT-5.6 Sol's promotional price lasts at least through 21 November 2026. Existing pins and bare aliases are preserved.

Correctness rules remain in force

Quota and price never lower a hard floor. Mini/high remains the bounded pipeline exception. Anthropic overrides retain Haiku 4.5, Sonnet 5 and Opus 5; Opus 5.5 adds a 20% input/output price saving, with provisional local fit. Sonnet/high is the only declared provider fallback for Terra/medium and Sol/medium; none is established for Sol/xhigh.

What remains provisional

OpenAI reports fewer factual mistakes for Sol and DeepSWE v1.1 scores of 68.8% for Sol and 66.6% for Luna at API max. Artificial Analysis v4.3.2 shows 48 and 37 in the September 25 snapshot. API max results do not measure our Codex medium/xhigh routes.

OpenAI announcement · Artificial Analysis Sol · Artificial Analysis Luna

Collect 80 completed cards across four task strata, paired deterministic runs, review grades, semantic reissues, critical defects, tokens, latency, and retry-adjusted cost. Historical GPT-5.6 results remain historical; GPT-6 outcome cost is unknown.

Both GPT-6 migrations remain proposal-only. An ultra source pin would become xhigh, an explicit reasoning downgrade. A controlled medium pilot cannot certify the xhigh correctness floor, a missing completion cohort, or runtime availability.