Forecast every task as a share of the five-hour cap.
Before launch, compare every selectable candidate model: median next-attempt demand, estimated percentage of one full five-hour quota window, uncertainty, and current headroom. Pick a model—or wait—with evidence.
One metric, two estimates
The provider does not publish the denominator, and a prompt does not reveal the next attempt's token consumption. Both sides must be learned—and both carry uncertainty.
Demand by candidate
Prompt size + task type + similar closed attempts on this exact model and thinking level produce a median token estimate and range.
Effective window
Clean same-window token/percentage observations estimate an empirical percentage-points-per-token slope for this plan and model.
Combine distributions
Multiply attempt demand by the calibrated slope. Report a median estimate, an 80% prediction range, confidence, and evidence age.
What the operator would see
Every selectable, nonrestricted candidate gets a row before current availability is attached. Current quota stays separate from the full-window forecast, and missing evidence stays visible.
| Candidate | Five-hour forecast | 80% range | Calibration | Confidence | Current state |
|---|---|---|---|---|---|
| CLI A · balanced · medium | 4.2% | 2.7–6.8% | Example Pro · cal-v3 · 2026-07-10 UTC | Medium · quantized meter | 31% left · 2026-07-11 14:20 UTC |
| CLI B · economy · medium | 2.1% | 1.2–4.4% | Example Plus · cal-v2 · 2026-07-09 UTC | Low · wide residuals | 64% left · 2026-07-11 16:05 UTC |
| CLI A · new · high | — | — | No compatible calibration | Uncalibrated · no safe transfer | Available · reset separate |
Illustrative UI copy only. A value above 100% is not clamped. Weekly and other provider limits remain separate admission inputs.
The plan in five focused pages
Start with the semantic trap, then follow the evidence into a deliberately simple forecast and clear repository boundaries.
What “% of cap” means →
Usage-anchored and rolling windows, opaque limits, and the two unknowns.
Calibration without pretending →
Existing Studio observations, the missing capture, pair eligibility, glitches, and confidence.
Forecast v1 →
A next-attempt cohort heuristic using prompt size, task type, model-specific history, and prediction ranges.
Who owns what →
Pure TokenEconomy math, Studio data and decisions, public website method.
Six small cards →
Inputs, outputs, dependencies, validation order, and the operator GO gate.
Full Markdown plan ↗
The reviewable source of truth in the public repository.
Language matters
Say
“Estimated 4.2% of one full effective 5h window; 80% range 2.7–6.8%; plan/model calibration as of …”
Do not say
“The provider gives this model N tokens every five hours.” The snapshots cannot establish that literal claim.
When evidence is weak
Show low confidence or uncalibrated. Unknown is safer and more useful than false precision.