
You are not buying a model. You are buying a unit of work. Same support workload, same volume, and the monthly bill runs from $216 to $20,000.
Three questions decide the bill: what value does the product deliver, can the charge cover delivering it, and who caps the risk when usage runs? Five answers, in order:

The proof at one support volume: Intercom Fin prices the resolved outcome, Salesforce Agentforce prices the conversation, Qwen prices the token. The model did not move the price. The unit did.

A flat subscription cannot absorb inference cost that swings this hard, which is why usage-based and hybrid shapes now set the market. Four numbers size the shift.

Which hours does your product take out of the workday, and can you count them? That answer is your value metric, and monetization starts there, one step before the charge metric everyone reaches for first. AI products deliver three kinds of value, and each one supports a different charge:
You learn what those benefits are from user conversations, lots of them.
"Before any number, we ask three questions: how many hours does this save per case, how many error points does it remove, and which risks does it retire? If nobody can answer, there is nothing to price."
The review we run with every enterprise team before we quote.
Focus on results. Customers will describe daily usage; watch the outcomes they achieve or hope to achieve instead. Then watch for the double-pay trap: an agent that fails mid-task bills twice, once for the attempt and once for the human who finishes. Outcome-based vendors bill per resolved result, not per attempt.
Outcome pricing is not practical for every product. Where it is practical, it is a bet: the seller absorbs every failed attempt, so the model holds only while failures cost less than the premium the outcome commands. Price the outcome you can count, not the activity you can watch.
Every quote hides a charge metric: the unit of usage a business prices against. AI makes that unit unusually hard to pick, because the cost behind it swings with the task and with the model. The metric still has to align with your value metric while reliably covering those costs. Every meter prices a risk before it prices a number: at the cost-aligned end, the buyer accepts a loose link to business value; at the outcome end, the seller prices costs it has never run.
"We pass inference cost through and take our margin on the value above it. When models get cheaper, your bill should shrink. That is the point of publishing list prices."
The Qwen pricing team, on why we price per token.

A base fee recurs per seat or platform; a scaling fee tracks your charge metric. Between the two extremes sit five shapes, each balancing predictable revenue against room to grow.
Customer acquisition and growth comes first: for an early-stage or product-led motion, every fixed fee is friction at the front door. Repeatable revenue comes next: recurring fees make revenue forecastable, and when one bill swings from, say, $100K to $300K month over month, the model turns hybrid. Above both sits one question: is this feature the product, or a component of it?
"Price a component to be used, not to be profited from."
The Qwen pricing team, on component pricing.

Our rate card is public, every number of it. Qwen3-Max lists at $1.20 input and $6 output per million tokens under 32K context, and $3 and $15 at the top tier. Qwen3.8-Flash lists at $0.15 and $0.47 up to a million-token context. New accounts start with a million free tokens for 90 days.
Steady load changes the shape. Dedicated model units sell capacity by the hour: an eight-unit MU1 serving Qwen3.6-Plus in Singapore runs $88 an hour.
Open weights change it again, and this is the lever a closed rate card cannot hand you. Qwen3.8 ships under Apache 2.0, so our token price becomes your compute price inside your own perimeter. A line you rent turns into capacity you own. Your exit option becomes real, and that strengthens your hand on every other line on the sheet. A published vLLM recipe fits Qwen3.8-27B at NVFP4 precision into 24.6 GiB on one Blackwell graphics processing unit. The same recipe holds 6.6 million key-value cache tokens at a one-million-token context.

Two levers move the number without moving the model. Batch inference bills at 50% of the real-time price on Model Studio as on other major APIs, and an explicit cache hit bills at 10% of the input price, per Model Studio pricing, Sep 2, 2026.
The bill that ends a contract is rarely the bill you modeled. A good charge metric and a balanced pricing model absorb some of that risk, but surprise bills stay structural wherever usage is variable. The right guardrails depend on which patterns of use are most likely to run away.
In every case the fix is communication before the bill, not after: show customers how usage becomes cost before they overspend. We charge usage to cover cost, not to maximize profit, and that philosophy only survives contact with a variable bill if the guardrails above sit in the contract, not only the console. Four clauses do most of the work: unit pricing fixed in writing, and an explicit no-training clause covering inputs and outputs. Add an exit path to self-hosted open weights if terms stop making sense, and agree portability before signature. Console budgets and alerts from day one cover the rest.
Guardrails are not distrust. They are what make usage-based pricing signable.
Take a support agent handling 10,000 conversations a month, at 8,000 input and 2,000 output tokens per conversation. Run that one workload through the three meters:
Per token, Qwen3-Max tier 1: 80M input tokens at $1.20 per million, plus 20M output tokens at $6.00 per million, equals $216.
Per conversation, Salesforce Agentforce: 10,000 conversations at $2.00, equals $20,000.
Per outcome, Intercom Fin: 10,000 resolutions at $0.99, equals $9,900.

The quote stops being a number and becomes a decision.
The demo bills the average. Production bills the tail. These three numbers assume 10,000 identical conversations. Real traffic never behaves: lengths vary, retries double the failed ones, and context grows as teams bolt on memory and tools. Price the distribution you will actually run, not the flat line in the pilot deck.
If you want these three numbers for your own volume, run your workload through the same three meters with the team.
The token line is the one every team argues about, and it is only one of six. The six do not move together: as model prices fall, tokens and serving shrink while talent and governance grow, so the total holds even as the mix flips. This is the sheet we ask teams to keep.

Every line on the sheet needs an owner from day one. An unowned line bills itself later, at a worse rate.
Compute gets cheaper and buyers get savvier while your price stands still. The failure that ends companies is scale that outruns accounting: volume climbs while unit margin goes negative and nobody has counted. Reprice on a calendar, not on a complaint. Five signals should trigger a pass:
"When we change a price fundamentally, we roll it to new workloads first and keep every other change small, localized, and weekly. Running workflows keep their economics until we say otherwise."
The Qwen pricing team, on repricing without breaking production.

The repricing date is a product decision, not a finance ritual.
Qwen meets every lane on that sheet with one model family. Run it as open weights inside your own perimeter, or as managed APIs on Model Studio. Price per seat with Token Plan when teams need one allowance across modalities, and reserve dedicated model units by the hour when load turns steady.
If you want to talk through pricing your AI product, start the conversation with the team: one minute, no login. See more of the enterprise AI portfolio at qwen.ai.
About Qwen. Qwen gives enterprises an open route into production AI: one model family covering language, vision, audio, coding, embeddings, and agents, as open weights and managed APIs. Built on curated data and ongoing research, Qwen supports creating, testing, and running AI systems for decisions that matter, from self-serve developer access to dedicated inference.
Sources: Alibaba Group Qwen open-model download milestone, 3B+ downloads (Aug 2026); MIT Project NANDA, The GenAI Divide (2025), the 95% pilot figure; vendor pricing pages accessed Sep 2, 2026: Intercom Fin, Salesforce Agentforce, GitHub Copilot, ElevenLabs, and Clay; Alibaba Cloud Model Studio pricing, Token Plan, dedicated model unit, and model deployment pages (verified Sep 2, 2026); the vLLM deployment recipe for Qwen3.8-27B (verified Sep 2, 2026).
Johnny Mai - Director, Gen AI Product Strategy & GTM, Qwen.
Disclaimer: This post is provided for general information only. Figures from third-party research and third-party pricing pages (MIT Project NANDA 2025; vendor pricing pages accessed Sep 2, 2026) are quoted as published in those sources and are not independently verified by Qwen or Alibaba Cloud. Pricing, product and service availability, features, and program terms vary by region and may change over time. Nothing in this post constitutes professional, legal, or investment advice. © 2026 Qwen. All rights reserved.
5 posts | 0 followers
FollowCommunity Builder - July 1, 2026
Community Builder - June 23, 2026
Alibaba Cloud Community - January 14, 2026
Justin See - March 19, 2026
Justin See - March 20, 2026
Alibaba Cloud Community - October 16, 2025
5 posts | 0 followers
Follow
Alibaba Cloud Academy
Alibaba Cloud provides beginners and programmers with online course about cloud computing and big data certification including machine learning, Devops, big data analysis and networking.
Learn More
DevOps Solution
Accelerate software development and delivery by integrating DevOps with the cloud
Learn More
Backup and Archive Solution
Alibaba Cloud provides products and services to help you properly plan and execute data backup, massive data archiving, and storage-level disaster recovery.
Learn More
Hybrid Cloud Solution
Highly reliable and secure deployment solutions for enterprises to fully experience the unique benefits of the hybrid cloud
Learn More