
One retrieval change took a knowledge assistant from $301.40 to $631.40 a month. Same users, same model, same price per token. About one in five respondents say AI-related operating costs, including token costs, constrained their AI use (McKinsey, The State of AI in 2026, August 2026). The price was never the problem. The forecast was.

Forecast the token bill of materials, not the price per token.
In this guide
What a token is, and why words are the wrong unit for a forecast
The four prices on every pricing page, and the details that move a bill more than the headline rate
A four-line forecasting worksheet you can fill in from your own logs
Two worked forecasts that show which lever moves each bill
A cost-control checklist to hold the forecast after go-live
LLM token pricing means you pay for the text a model reads and writes, measured in tokens.
A token is a chunk of text the model's tokenizer produces: sometimes a whole word, often part of a word, sometimes a punctuation mark or a space.
For forecasting, the catch is that tokens are not words, and the ratio depends on language. For Qwen models, the rule of thumb in the official documentation is that "1 token is 3~4 characters for English texts and 1.5~1.8 characters for Chinese texts," with a vocabulary of 151,646 tokens (Qwen documentation, Key Concepts, accessed Oct 5, 2026).

So the same paragraph translated into Chinese can tokenize very differently from the English original. For a first-pass English estimate, I plan on roughly 1.3 tokens per word. That is my planning assumption, not a measured figure: replace it with counts from your own logs within the first week.
For a FinOps lead or engineering manager, this is a usage-based bill: predictable only if you can predict the usage. Most AI cost surprises are not price surprises. They are usage surprises measured in tokens nobody counted.
If your first workload is still measured in users rather than tokens, map one workload into a token bill of materials with the team, starting from the logs you already have.
Count tokens from logs, not words from documents.
Every serious LLM price list has four lines.
On Alibaba Cloud Model Studio, they look like this for the current Qwen models (Singapore list prices, per 1M tokens; Alibaba Cloud, Model Studio pricing, accessed Oct 5, 2026):

Four details change forecasts more than the headline prices do:
1. Output is the expensive side. On qwen3.8-max, one output token costs the same as three input tokens. A workload that writes long answers or reasons at length is an output-heavy workload, and its bill moves with output length.
2. Batch is a different model ID, not a switch. In Singapore, batch runs on the models the batch page lists. For example, qwen-flash lists at $0.05 input and $0.40 output per 1M tokens (up to 256K input), so batch brings it to $0.025 and $0.20 (Alibaba Cloud, Batch inference and Model Studio pricing, accessed Oct 5, 2026).
3. Discounts do not stack. Batch and context cache discounts cannot apply simultaneously (Alibaba Cloud, Model Studio pricing, accessed Oct 5, 2026). Forecast one discount per workload, not both.
4. Some models are tiered by input length. qwen3-max, for example, is priced from $1.20 input and $6.00 output per 1M tokens for inputs up to 32K tokens, rising at longer inputs. If your prompts get longer, the price per token can change, not only the token count.
The price page has four lines; your forecast needs all four.
The forecasting unit is not the token and not the user.
It is the token bill of materials: for one request, what goes in, what comes out, and at which price. Multiply it by how many requests you expect, and you have a forecast you can defend.
The worksheet has four lines:

Build the input tokens per request from its parts, because each part grows differently:
System prompt and instructions. Fixed per request. Often a few hundred to a few thousand tokens.
Retrieved context. Documents pulled in by retrieval. Scales with how many chunks you retrieve and how long they are.
Conversation history. In multi-turn chat, earlier turns are usually re-sent, so input per request grows with conversation length.
User message. Usually the smallest part.
Output has two parts: the visible answer and, on thinking-mode models, the reasoning the model generates before it. For the forecast, treat reasoning output as billable output and confirm the billing per model and mode on the pricing page.
Write the bill of materials before you write the forecast.
Two workloads, same pricing page, very different bills.
All prices are Singapore list prices per 1M tokens (Alibaba Cloud, Model Studio pricing, accessed Oct 5, 2026). Volumes and token counts are illustrative assumptions.
Workload A: internal knowledge assistant on qwen3.8-flash ($0.15 input, $0.47 output).
Assumptions: 2,000 employees, 10 requests per person per working day, 22 working days. Each request sends 3,000 input tokens (system prompt, three retrieved passages, the question) and returns 500 output tokens.

Result: $301.40 per month, or $3,616.80 per year.
Sensitivity test. The team retrieves more context, and input per request rises from 3,000 to 8,000 tokens. Input becomes 440,000 x 8,000 = 3,520M tokens x $0.15 = $528.00. Total: $528.00 + $103.40 = $631.40 per month.
Usage did not change. The bill more than doubled, driven by one line of the bill of materials.
Workload B: contract review on qwen3.8-max ($2.00 input, $6.00 output).
Assumptions: 20,000 documents per month, each sending 40,000 input tokens (the contract plus a review rubric) and returning 4,000 output tokens (findings plus reasoning).
Contract review rarely needs an answer in seconds. If your evaluation shows the batch-eligible qwen-max (Singapore list $1.60 input, $6.40 output per 1M tokens) meets the quality bar, batch at 50% gives $0.80 input and $3.20 output:

Result: $2,080.00 per month in real time, or $896.00 per month on batch, 57% less than the real-time forecast.
Two lessons from the same page. In Workload A, input context is the lever. In Workload B, latency tolerance is the lever.
Neither lesson is visible from the price per token.
Find the one line that doubles your bill
Send request volumes and token counts per workload, from last month's logs or honest estimates. The team will price each bill-of-materials line on the tier and mode that fits, and show which lever, context or latency, moves your number most. Price my workloads line by line →
Forecast each workload on its own lever, then add them up.
These are the mistakes that turn a reasonable pilot into an awkward budget review.
1. Forecasting in words, not tokens. A words-to-tokens guess is fine for day one. After that, every forecast should use token counts from your own request logs, split by language. A bilingual workload can tokenize very differently from an English one.
2. Forgetting that input grows. Retrieved context lengthens as the knowledge base grows, chat history is re-sent each turn, and agents feed tool results back as input. Context design matters enough that Alibaba Cloud says its new Agent Context layer cuts "token usage by up to 67%" (Alibaba Cloud, Apsara Conference 2026, September 2026). If context can move token usage that much, your forecast has to model context, not just traffic.
3. Ignoring output and reasoning length. Output is the expensive side of every price list. A prompt change that makes answers twice as long, or switches on step-by-step reasoning, can move the bill more than a traffic spike. Set and track a maximum output length per workload.
4. Forecasting the average and missing the tail. Retries, timeouts, loops and the 1% of requests with enormous documents do not show up in an average. Forecast the 95th percentile request as well as the mean, and put retries in the model explicitly.
5. Pricing at the wrong line. Using the wrong region, the wrong tier, a discount that does not apply to that model, or two discounts that do not stack. Free quota belongs here too: new Model Studio users in Singapore get 1M tokens per model for 90 days (Alibaba Cloud, Model Studio pricing, accessed Oct 5, 2026). That is evaluation budget, not a production forecast.

Every forecasting mistake is a line missing from the bill of materials.
A forecast is a promise; controls are how you keep it.
Use this checklist per workload:
1. Tag every request to a workload and owner. Unattributed tokens cannot be managed. Model Studio provides call statistics and cost analysis in the console; for code-first agents, AgentRun advertises "token-level cost attribution" (Alibaba Cloud, What is AgentRun, accessed Oct 5, 2026).
2. Route to the smallest tier that passes your evaluation. qwen3.8-flash input costs $0.15 per 1M tokens against $2.00 on qwen3.8-max. Send work to the flagship only where your evaluation shows it is needed.
3. Cap output. Set a maximum output length per workload and alert when real output approaches it.
4. Trim the input. Retrieve fewer, better chunks. Summarize long histories. Remove instructions nobody reads.
5. Move non-urgent work to batch. Reports, evaluations, enrichment and document review can usually wait hours. Batch is 50% of real-time price on supported models.
6. Keep stable prefixes stable. Put fixed instructions at the start of the prompt so repeated content is eligible for cache hits. Measure the hit rate rather than assuming one.
7. Set budgets and alerts by workload, with a soft limit that notifies and a hard limit that requires approval to raise.
8. Reforecast monthly. Compare forecast to actual per bill-of-materials line, and fix the line that drifted.

When not to forecast purely per token. Per-token pricing is the right default for new and variable workloads. It is not always the right contract. If a workload's usage is steady and high, look at the alternatives Model Studio lists: savings plans and Throughput Reservation, which locks dedicated inference throughput for a model, for predictable volume, and the Coding Plan at a fixed monthly fee for AI coding tools (Alibaba Cloud, What is Model Studio and Throughput Reservation, accessed Oct 5, 2026).
For individual developers, Qwen Cloud lists an Individual Token Plan from $6 per month (Qwen Cloud, accessed Oct 5, 2026). Forecast in tokens either way. Then choose the contract the forecast supports.
Controls turn a forecast into a budget you can hold.
You can produce a defensible forecast in five working days.
1. Day 1: List workloads. Every place your organization calls a model, with an owner.
2. Day 2: Write the bill of materials. For each workload, input tokens per request by part (system, retrieval, history, user) and output tokens per request.
3. Day 3: Pull real counts. Replace estimates with token counts from logs or a short test run. Split by language.
4. Day 4: Price each line. Exact model ID, region, tier and the one discount that applies. Run the 95th percentile request as well as the mean.
5. Day 5: Set controls. Owner tags, output caps, a budget alert and a reforecast date.

For a FinOps lead, this matters beyond the AI line item. Boston Consulting Group reports that AI spending has grown from about 1.7% of revenue in late 2025 to 3.3%, with more than 80% of it sitting outside enterprise IT (BCG, Applied AI Index 2026, September 2026). Spend that lives in business budgets needs a forecast business owners can read, the same ownership discipline behind the build-or-partner call in Build vs. Buy, or Both? The Enterprise AI Decision.
A token bill of materials is that forecast.
Forecast the unit you can count, then hold every workload to it.
Key takeaways
Forecast the bill of materials. Input tokens, output tokens and the price of each, multiplied by volume, beat any price-per-token comparison.
Find each workload's lever. One retrieval change moved Workload A from $301.40 to $631.40; batch took Workload B from $2,080.00 to $896.00.
Control after go-live. Owner tags, output caps, budget alerts and a monthly reforecast keep the number you promised.
Walk into budget review with a forecast finance can read
List your workloads, owners and rough volumes in a short form, and the team will walk through the four-line worksheet with you. Every line gets checked for region, tier and the one discount that applies, and for whether a savings plan or batch fits better than pay per token. Pressure-test my token forecast →
See more of the enterprise AI portfolio at qwen.ai. Lasting returns come from treating AI like infrastructure: planned over years, delivered in stages, with Qwen alongside you so the early choices hold up and the value keeps growing.
About Qwen. Qwen gives enterprises an open route into production AI: one model family covering language, vision, audio, coding, embeddings, and agents, as open weights and managed APIs. Built on curated data and ongoing research, Qwen supports creating, testing, and running AI systems for decisions that matter, from self-serve developer access to dedicated inference.
Sources: McKinsey, The State of AI in 2026: On the Road to ROI (August 2026); BCG, Applied AI Index 2026 (September 2026); Qwen documentation, Key Concepts (accessed Oct 5, 2026); Alibaba Cloud, Model Studio pricing (accessed Oct 5, 2026); Alibaba Cloud, Context cache (accessed Oct 5, 2026); Alibaba Cloud, Batch inference (accessed Oct 5, 2026); Alibaba Cloud, What is AgentRun (accessed Oct 5, 2026); Alibaba Cloud, What is Model Studio (accessed Oct 5, 2026); Alibaba Cloud, Throughput Reservation (accessed Oct 5, 2026); Alibaba Cloud blog, Apsara Conference 2026 (accessed Oct 5, 2026); Qwen Cloud (accessed Oct 5, 2026).
Johnny Mai - Director, Gen AI Product Strategy & GTM, Qwen.
Disclaimer: This post is provided for general information only. Worked examples are illustrative and use stated assumptions; your costs will vary. Prices are Alibaba Cloud Model Studio list prices for the Singapore region as published in October 2026, vary by region, and may change over time. Figures from third-party research (McKinsey 2026; BCG 2026) are quoted as published in those reports and are not independently verified by Qwen or Alibaba Cloud. Product and service availability, features, and program terms vary by region and may change over time. Nothing in this post constitutes professional, legal, or investment advice. © 2026 Qwen. All rights reserved.
Model Routing Explained: Send Each Request to the Cheapest Model That Passes
Should This Workflow Be an Agent? A Three-Question Test Before You Build
14 posts | 2 followers
FollowJohnny Mai - October 7, 2026
Johnny Mai - September 10, 2026
Farruh - May 26, 2026
Farruh - June 9, 2026
Alibaba Cloud Community - August 5, 2026
Hosung Kim - July 21, 2026
14 posts | 2 followers
Follow
Qwen
Full-range, open-source, multimodal, and multi-functional
Learn More
Alibaba Cloud Model Studio
A one-stop generative AI platform to build intelligent applications that understand your business, based on Qwen model series such as Qwen-Max and other popular models
Learn More
Token Plan
Build more, spend less. One plan, every modality.
Learn More
Alibaba Cloud for Generative AI
Accelerate innovation with generative AI to create new business success
Learn MoreMore Posts by Johnny Mai