Start with a cohort heuristic we can explain and falsify.
For each candidate model and thinking level, find genuinely similar closed attempts, estimate a token distribution, and combine it with that scope's calibrated meter slope.
Three inputs per candidate
Prompt size
A stable measure at the forecast boundary, such as UTF-8 bytes, with versioned rules for wrappers and attachments. It is a cheap proximity feature—not a tokenizer promise.
Task type
Map Studio task metadata onto TokenEconomy's existing task classes where possible: design, feature, chore, doc edit, research.
Similar closed attempts
Exact model + thinking level + task type + compatible run-policy version, preferring the nearest prompt-size band. Keep outcome, censoring, repository, and timestamps.
What v1 forecasts
One next launch attempt
Consumption from run start through terminal state under the current failure/cancellation policy. Successful, failed, and cancelled attempts count when tokens are complete.
Not eventual successful completion
Retries, reissues, and review loops are not included. That larger completion-cost question needs a separate future model.
Cap walls are lower bounds
Quota-truncated or still-running attempts are right-censored. V1 reports their rate and lowers confidence instead of inserting the observed partial total as exact demand.
The v1 recipe
Repeat the same deterministic procedure for every selectable, nonrestricted model and proposed effort before attaching current CLI availability. Do not estimate one universal attempt total and reuse it across models.
Uncalibrated. A later fallback hierarchy requires replay-validated, frozen transfer rules and explicit interval widening.central = median(draws) · range80 = [q10(draws), q90(draws)]
Result contract: every candidate gets a row
The result is more than a number. Scope, age, uncertainty, evidence, and failure status travel with it.
| Candidate | 5h forecast | 80% range | Confidence | Current quota | Median attempt tokens | Calibration | Demand evidence |
|---|---|---|---|---|---|---|---|
| CLI A · balanced · medium | 4.2% | 2.7–6.8% | Medium · 7 windows · quantized meter | 31% left · resets 2026-07-11 14:20 UTC | 186k | Example Pro · 5h · cal-v3 · as of 2026-07-10 UTC | effective n=42 · mixed outcomes · 5% censored · policy-v4 |
| CLI B · economy · medium | 2.1% | 1.2–4.4% | Low · 3 windows · wide residuals | 64% left · resets 2026-07-11 16:05 UTC | 214k | Example Plus · 5h · cal-v2 · as of 2026-07-09 UTC | effective n=19 · mixed outcomes · 16% censored · policy-v4 |
| CLI A · new · high | — | — | Uncalibrated · no safe transfer | Available · reset separate | — | No compatible calibration | No compatible cohort |
Illustrative values only. The current quota column is supplied separately by Studio and is not folded into the full-window forecast.
Uncertainty and display rules
Prediction, not confidence, range
The attempt range includes run-to-run variation and calibration uncertainty. A separate confidence label reports evidence quality.
No false decimals
Retain precision in the API; normally show at most one decimal and respect the source meter's quantization. Include scope and as-of time.
Do not clamp
Forecasts may exceed 100%. Unknown stays unknown. Neither case becomes zero or silently disappears from model comparison.
SuggestModel suitability beside this metric; never rank candidates by lowest cap percentage alone.Named upgrade paths
- V2 per-repository/per-task-type regression: prompt size, repository, model, thinking, task type, and interactions, evaluated on future time periods.
- Matched or controlled crossover evidence to measure historical model-routing bias.
- Backtested hierarchical pooling for sparse plans/models, with transfer error reflected in wider ranges.
- Separate input/output/cache/reasoning slopes when the data can identify them.
- Censoring/survival treatment for interrupted work and richer features such as changed-file scope, gates, or tool intensity.