Input, output, and cache-hit token totals are multiplied by their applicable coefficients and added for each request.
The same count can imply different point use when Model ID, token mix, coefficient table, or input band changes.
Its prose says 12,985 points; the displayed factors and arithmetic produce 26,475 on the page checked 4 August.
Budget language tends to harden before anyone asks what it measures. A team predicts thirty million tokens, attaches a currency value or plan tier, and carries the number through procurement as if every token were interchangeable. That assumption can survive a model change, a longer context window, a new cache strategy, and an altered response ratio without ever appearing to change. The spreadsheet remains calm because it has discarded the mechanism beneath the total.
TokenHub’s revised personal General Token Plan makes that discarded mechanism explicit. Points are the plan’s common consumption unit. Tokens remain the observed quantities, but they are sorted into input, output, and cache-hit classes, then weighted. The budget number is no longer the count. It is the result of a rule applied to the count.
A 1:1 allowance conversion is not a permanent consumption identity
The official announcement says the personal General Token Plan changed its measurement unit from tokens to points from 20 July 2026 Beijing time. Existing plan allowances were converted 1:1: a numerical allowance denominated in tokens became the same numerical allowance denominated in points. The same announcement then gives separate point-deduction coefficients for input, output, and cache-hit tokens. source ↓
The distinction is elementary and easy to erase. A conversion describes what happened to the balance at the transition. A deduction rule describes what future requests remove from that balance. Converting 100 tokens of old allowance into 100 points does not promise that every later token consumes one point. Some listed existing models retain coefficients of 1.00 across all three classes; models listed as arriving on or after 20 July use differentiated coefficients.
Editorial judgment: any document that summarizes the change as “tokens were simply renamed points” is operationally defective. It preserves the opening balance while deleting the new consumption function. The plan may display one common unit, but the path into that unit is model- and workload-dependent.
The formula has three token classes, not one total
The official rules define per-request consumption as input tokens multiplied by the input coefficient, plus output tokens multiplied by the output coefficient, plus cache-hit tokens multiplied by the cache coefficient. The rules page reports its own update timestamp as 14 July 2026 10:35:00. source ↓
That structure destroys the sufficiency of a single token total. A workload with heavy output can consume differently from one with the same total dominated by cache hits. A cache policy is therefore not only a latency or infrastructure detail inside this plan; its measured hit volume enters the point equation with its own coefficient. The input-to-output ratio is also budget state.
Forecasting from “tokens per month” now omits at least three pieces of information: which portion is input, which is output, and which input is counted as cache-hit use under the rule. Even before model selection changes, two workloads with the same aggregate total can land on different weighted sums.
The exact Model ID is part of the budget contract
The rules table makes names insufficient. It lists two DeepSeek-V4-Flash IDs that use the same underlying scheduled model but different deduction methods. deepseek-v4-flash-202605 carries coefficients of 1.00 input, 1.00 output, and 1.00 cache-hit. deepseek-v4-flash-202607 carries 2.18 input, 4.35 output, and 0.05 cache-hit. The corresponding Pro entries also divide: deepseek-v4-pro-202606 is listed at 1.00/1.00/1.00, while deepseek-v4-pro-202607 is listed at 6.53/13.05/0.06.
This is more than price-table trivia. A budget labeled only “DeepSeek-V4-Flash” cannot reveal which deduction mapping was applied. Two calls can reach the same underlying scheduled model while crossing the plan ledger differently because their Model IDs differ. The identifier is not decorative metadata. Under this rule it selects economic state.
MiniMax-M3 adds a second axis. For inputs at or below 512k, the table lists 2.29 input, 9.14 output, and 0.46 cache-hit. Above 512k, it lists 4.57, 18.27, and 0.92. A forecast that records the Model ID but drops the applicable input-length band is still incomplete.
A family label cannot distinguish the older and newer DeepSeek deduction mappings.
MiniMax-M3 changes coefficients when input moves above 512k.
The same identifier and workload need the dated coefficients actually applied in that window.
Reference estimates are not workload receipts
The official rules publish estimated usable-token figures for plan tiers, then bound them carefully. They are references, not promises of actual usable tokens. The page says actual use depends on the workload’s cache hit rate, input/output token ratio, model mix, and real-time pricing rules. Those are precisely the variables a one-line token budget tends to remove.
Inference: a procurement model built from the published estimate can be directionally useful, but it cannot be reconciled after the fact unless the team also measures its own mix. The estimate is a scenario. The workload receipt is evidence. Confusing them makes variance look mysterious when the omitted variables were named in advance.
This issue does not convert points into an API price, predict an invoice, or extend the personal-plan rule to enterprise arrangements. The checked sources do not establish those claims. They also do not guarantee that today’s coefficients will remain fixed. A versioned table matters because the rules themselves identify a dated effective state.
Example 2 must remain unresolved
The official rules page contains an internal contradiction. Example 2 describes 10,000 input tokens, 500 output tokens, and 50,000 cache-hit tokens for deepseek-v4-flash-202607. Its prose says the request consumes 12,985 points. The displayed calculation uses coefficients 2.18, 4.35, and 0.05:
10,000 × 2.18 + 500 × 4.35 + 50,000 × 0.05 = 26,475
Unknown: as checked on 4 August 2026, the public page does not reconcile 12,985 with 26,475. The arithmetic shown and its stated factors produce 26,475; the prose states 12,985. This issue will not silently choose a corrected official value, infer which field is mistaken, or assign intent. Operators should treat the example’s result as unresolved and obtain authoritative clarification before using it as an acceptance fixture.
Keep the rule that made the number
Editorial recommendation: retain a versioned billing receipt for each workload window. It should contain the exact Model ID; input, output, and cache-hit token totals; the applied input, output, and cache coefficients; any input-length band; the plan and rule effective date; and the computed point consumption. Preserve the source or immutable version identifier for the coefficient mapping, not merely a screenshot of the final balance.
The receipt should be generated near the measurement boundary, then joined to procurement forecasts by stable workload and period identifiers. If a gateway rewrites a model alias into an exact ID, preserve both requested and resolved identifiers. If cache accounting is supplied by the service, distinguish reported cache-hit tokens from locally estimated hits. If the computation is aggregated, retain enough per-model and per-band detail to reproduce the total.
Alert when an exact Model ID maps to a different coefficient set, when a new ID enters the workload, when an input crosses a coefficient threshold, or when the effective date changes. The alert should block silent continuation of the old budget assumption until the forecast is recomputed against a measured mix. A changed table is not ordinary spend variance; it changes the function that defines spend inside the plan.
- 01Resolve the workload identity.
Record requested and exact resolved Model IDs, plus the plan scope to which the deduction rule applies.
- 02Measure the three token classes.
Keep input, output, and cache-hit totals separate; do not reconstruct them from one aggregate count.
- 03Bind the dated coefficients.
Store each applied coefficient, the input band, rule effective date, and a versioned reference to the mapping.
- 04Recompute and reconcile.
Calculate points from the preserved fields, compare with the authoritative plan record, and leave contradictions visible.
- 05Alert on rule drift.
Require budget review when identifiers, mappings, coefficients, thresholds, or effective dates change.
Limit
This is a source-bounded reading of the official TokenHub announcement and personal General Token Plan rules as checked on 4 August 2026. Receipt fields, monitoring, reconciliation, and acceptance gates are editorial engineering recommendations—not Tencent Cloud policy, accounting advice, or a statement about invoice settlement.
What would change this assessment
A revised official rule that makes one unweighted token total sufficient; an authoritative correction reconciling Example 2; a versioned machine-readable coefficient endpoint with historical mappings; or plan records that expose the exact per-window inputs, coefficients, and computed points needed for independent reconciliation. Each would change a specific operational burden. None is established by the two pages checked here.
A count without its rule is not a budget. It is a number waiting to betray one.