¢ Token Economy
Research plan · no implementation

Forecast every task as a share of the five-hour cap.

Before launch, compare every selectable candidate model: median next-attempt demand, estimated percentage of one full five-hour quota window, uncertainty, and current headroom. Pick a model—or wait—with evidence.

Operator gate: this is the TE-4 concept deliverable. Every proposed implementation slice remains blocked until an explicit GO. Example values below are illustrative, not measured provider data.

One metric, two estimates

The provider does not publish the denominator, and a prompt does not reveal the next attempt's token consumption. Both sides must be learned—and both carry uncertainty.

01 · ATTEMPT

Demand by candidate

Prompt size + task type + similar closed attempts on this exact model and thinking level produce a median token estimate and range.

02 · METER

Effective window

Clean same-window token/percentage observations estimate an empirical percentage-points-per-token slope for this plan and model.

03 · FORECAST

Combine distributions

Multiply attempt demand by the calibrated slope. Report a median estimate, an 80% prediction range, confidence, and evidence age.

cap % draws = slope draws × attempt-token draws · show median and 10th–90th percentiles

What the operator would see

Every selectable, nonrestricted candidate gets a row before current availability is attached. Current quota stays separate from the full-window forecast, and missing evidence stays visible.

CandidateFive-hour forecast80% rangeCalibrationConfidenceCurrent state
CLI A · balanced · medium4.2%2.7–6.8%Example Pro · cal-v3 · 2026-07-10 UTCMedium · quantized meter31% left · 2026-07-11 14:20 UTC
CLI B · economy · medium2.1%1.2–4.4%Example Plus · cal-v2 · 2026-07-09 UTCLow · wide residuals64% left · 2026-07-11 16:05 UTC
CLI A · new · highNo compatible calibrationUncalibrated · no safe transferAvailable · reset separate

Illustrative UI copy only. A value above 100% is not clamped. Weekly and other provider limits remain separate admission inputs.

Language matters

Say

“Estimated 4.2% of one full effective 5h window; 80% range 2.7–6.8%; plan/model calibration as of …”

Do not say

“The provider gives this model N tokens every five hours.” The snapshots cannot establish that literal claim.

When evidence is weak

Show low confidence or uncalibrated. Unknown is safer and more useful than false precision.