
Paste a 500,000-token policy manual into every prompt: 200,000 monthly answers on qwen3.7-plus cost $120,792. Retrieve the five passages that matter: about $467. Re-embed the whole manual after a policy change: under four cents. Fine-tuning would have fixed none of it.

Retrieval is how a model learns what is true today; fine-tuning is how it learns how you work. Buy each for the job it does.
In this guide
What each tool actually changes: behavior per call, knowledge at answer time, or the weights
A four-question test that names the defect before you fund the fix
One workload priced four ways, with every assumption stated
Four patterns for combining RAG and fine-tuning in production
The cases where fine-tuning is the wrong tool, and a five-step start for this week
Most teams ask "RAG or fine-tuning?" as if the two compete.
They do not. They change different things, and the cost of choosing wrong is not a slightly worse model: it is a project that solves the wrong problem.
The pressure to get it right is real. 89% of respondents report regular AI use in at least one business function, yet 37% say AI has contributed positively to their EBIT (McKinsey, The State of AI in 2026: On the Road to ROI, Aug 2026). The gap between use and impact is filled with pilots that picked an expensive fix for a cheap problem.

AI engineering leads face two bad options. Fine-tune early and you own a training pipeline, a dedicated deployment and a retraining calendar before you know whether the model needed any of it. Stay on prompts forever and you pay for the same 1,500-token instruction block on every call, with outputs that still drift from your format.
The short answer: prompting changes behavior per call. Retrieval-augmented generation (RAG) changes what the model knows at answer time by fetching relevant passages from your documents into the prompt. Fine-tuning changes the model's default behavior and format by training its weights on your examples.
Rent knowledge, through retrieval, because knowledge changes. Own behavior, through fine-tuning, only when it is stable, proven and paid for.
Holding a pile of failing answers right now? Check whether your gap is knowledge or behavior with the team before you budget a fix.
Name the defect before you name the fix.
The comparison below is the one to put in your design review.
Prices and data minimums come from the Alibaba Cloud Model Studio documentation (accessed Oct 5, 2026); time estimates are the author's operating experience, not a benchmark.

Read the table by rows, not columns. If your problem sits in the "what it changes" row under knowledge, no amount of training fixes it reliably. If it sits under format at volume, no amount of retrieval does.
Knowledge belongs in an index you can update; behavior belongs in the weights only when it has stopped changing.
Answer these in order, and stop at the first decisive answer.
1. Does the paste test fix it? Take ten failing cases and paste the correct source passage into each prompt by hand. If the answers become right, the gap is knowledge, and you need RAG. If they stay wrong with the right facts in front of the model, the gap is behavior.
2. Does the knowledge change faster than you could retrain? Prices, policies, inventory, product specs, contracts: if any of it changes monthly, it belongs in retrieval. A fine-tuned fact is a snapshot with no expiry date printed on it.
3. Have prompts plateaued on a golden set, and do you hold 1,000+ quality examples? Build a scored evaluation set first. If careful prompting and few-shot examples stop improving the score, and you have at least 1,000 clean examples of the target behavior (the Model Studio minimum guidance for SFT), fine-tuning becomes a candidate. Without both, it is a guess with a training bill.
4. Does the volume pay back training and hosting? Fine-tuning earns its cost by deleting prompt tokens, moving work to a smaller model, or raising pass rates enough to cut retries and human review. Put a monthly dollar figure on that saving and compare it with training runs, dedicated deployment and data labeling over a year.

Most workloads end at question 1 or 2. The ones that reach question 4 with a yes are worth fine-tuning, and usually worth fine-tuning on top of RAG, not instead of it.
Prompt first, retrieve second, train last, and only with a score to beat.
One illustrative workload, an internal policy assistant, with every assumption stated.
Assumptions.
Volume: 200,000 questions per month.
Corpus: policy corpus of 2,000 pages at 250 tokens per page: 500,000 tokens.
Request shape: system prompt 1,500 tokens, question 200 tokens, answer 400 tokens.
Generator: qwen3.7-plus at $0.40 input / $1.60 output per 1M tokens for inputs up to 256K, and $1.20 / $4.80 above 256K (Model Studio, Singapore list, accessed Oct 5, 2026).
Retrieval: RAG retrieves five chunks of 500 tokens each.
Embeddings: text-embedding-v4 at $0.07 per 1M tokens (Singapore).
Discounts: no cache or batch discounts.

Result: $264 a month for prompting alone, $120,792 for the whole manual in every prompt, and about $467 for RAG.
Option A is cheap, and wrong whenever the answer lives in the manual. Option B shows that long context is a capability, not a knowledge strategy.
RAG costs about $203 a month more than prompting alone ($200 in generator tokens plus $2.80 in query embeddings), and it is the only option that answers from today's manual.
Option D: fine-tuning. Model Studio bills training as (training tokens + mixed training tokens) x epochs x training unit price. In Singapore, SFT on qwen3.8-27b lists at $0.0075 per 1K training tokens ($7.50 per 1M) and on qwen3-14b at $0.0016 per 1K ($1.60 per 1M) (Model Studio model training documentation, accessed Oct 5, 2026). Assume 2,000 examples of 800 tokens, 3 epochs, no mixed data:

Result: one training run costs $36 on qwen3.8-27b or $7.68 on qwen3-14b. The expensive lines sit around it.
The training run is the cheap line. A fine-tuned model must be exported as a snapshot and deployed before use, and the documentation itself says "the model deployment billing is high"; check your console for the deployment rate, which this article does not quote. Then add the human hours to write and review 2,000 examples, usually the largest cost of all, and one retraining run per material policy change and per base model you want to adopt.
What could fine-tuning save here? Training the system prompt into the weights would delete 1,500 input tokens per call: 200,000 x 1,500 x $0.40 / 1,000,000 = $120 per month on qwen3.7-plus. And it would not teach the manual, so the roughly $203 a month for retrieval stays.
A fine-tune for this workload has to beat $120 a month after training runs, dedicated deployment and labeling. At this volume it does not come close. At 100 times the volume, with a stable format, it might.
See all three options priced on your workload
Send your volume, corpus size, refresh rate and the defect you are trying to fix. The team will price prompt, RAG and fine-tune side by side on your numbers, including the deployment and labeling lines a training quote leaves out. Price my three options side by side →
Compare monthly cost per correct answer, not the sticker price of each technique.
The strongest production systems use more than one tool, each on the job it does best.
Four patterns cover most of them.
1. Prompt plus RAG. The default for knowledge work: a stable instruction block, retrieved passages, a fixed output schema. Start here for assistants, search and support.
2. RAG plus a light fine-tune. Retrieval supplies the facts; a LoRA fine-tune on a 9B to 27B model teaches the behavior around them: cite the passage, follow the house answer format, say "not in the sources" instead of guessing. Knowledge stays updatable; behavior stops drifting. Train on examples that include retrieved context, so the model learns to use passages rather than memorize them.
3. A fine-tuned small model in front, a large model behind. A fine-tuned 9B or 14B model handles a narrow, high-volume step, such as classification, extraction or query rewriting, and passes the hard cases to a larger managed model. This is where fine-tuning most often pays: narrow task, stable labels, high volume.
4. Continual pre-training plus RAG. For domains with heavy specialist vocabulary, continual pre-training on 50 million+ tokens of domain text helps the model read your documents; retrieval still supplies the current facts. Check region support first: Singapore currently supports SFT only, not continual pre-training or DPO.

Owning the retrieval pipeline versus running it on a managed service is the same build-or-partner call we broke down in Build vs. Buy, or Both? The Enterprise AI Decision: own the data and the evaluation set, and partner for the parts that are not your edge.
Combine tools by layer, facts from the index and habits from the weights, never the same job twice.
Fine-tuning is the right tool for fewer problems than its reputation suggests.
Skip it when:
The problem is facts. Stale prices, new policies, private records. A model trained on them will repeat them confidently after they change, with no source to check.
You cannot write the eval set. Without a scored golden set, you cannot show the fine-tune beat the prompt. You will ship it on a feeling.
You hold fewer than 1,000 clean examples. Below the documented SFT guidance, effort goes into data creation, not model improvement. Spend it on better prompts and retrieval first.
You need citations or an audit trail. Reviewers and regulators ask where an answer came from. Retrieved passages can answer that; weights cannot.
The prompt has not been tried properly. Few-shot examples, a strict schema and a clear instruction block solve most format problems in hours.
Volume is low. At the 200,000-question volume above, the most a fine-tune could save is $120 a month. Training, hosting and labeling will cost more.
Your base model will move soon. A fine-tune is pinned to the snapshot you trained on. Every base upgrade you want to adopt is another training run.

If retraining is how you fix a typo in a fact, you picked the wrong tool.
Five steps, each earned by a score, get you to the cheapest fix that passes.
1. Build the golden set. Collect 100 to 300 real questions with correct answers. Score every option against it, starting with Model Studio's Model Evaluation, which supports LLM-as-judge scoring, string match, text similarity and human annotation (Model Studio Model Evaluation documentation, accessed Oct 5, 2026).
2. Run the paste test. Ten failing cases, correct passages pasted in by hand. Knowledge or behavior: now you know.
3. Fix the prompt. Instruction block, few-shot examples, output schema. Re-score.
4. Add retrieval if the paste test said knowledge. In Model Studio, the managed Knowledge Base is free to create and manage, with hybrid retrieval and the qwen3-rerank reranker; you pay for the extra input tokens that retrieved chunks add. Note the access rule: the documentation states that only users who created Model Studio applications in the Singapore region before April 21, 2025 can access the Application Development tab, where the Knowledge Base sits, or call the knowledge base APIs. Other international teams build the same pipeline through the API with text-embedding-v4 and qwen3-rerank, or talk to the team about options (Model Studio knowledge base documentation, accessed Oct 5, 2026).
5. Fine-tune only with a score to beat. If prompts plus retrieval plateau, and question 4 of the test says the volume pays, train a LoRA SFT on a Singapore-trainable model with your 1,000+ examples and compare it on the same golden set. Model Studio states that it "will never use your data for model training" (Model Studio privacy notice, accessed Oct 5, 2026).

One model family, three tools, one evaluation set: the cheapest fix that passes wins, and the evaluation set decides.
Earn each layer with a score, not a slide.
Key takeaways
Name the defect first. Wrong facts call for retrieval; wrong behavior or format calls for prompting, and only later for fine-tuning.
Price per correct answer. On the same 200,000 questions, RAG costs about $467 a month against $120,792 for the whole manual in every prompt.
Train last, with a score to beat. Fine-tune only with 1,000+ clean examples, a plateaued golden set and a volume that pays back training and hosting.
Bring ten failing cases; leave knowing which fix to fund
Share the failing cases and your golden set, or your plan to build one, through a short form. The team will walk through the four-question test with you and map which layer, prompt, retrieval or weights, earns the next sprint. Walk my failing cases through the test →
See more of the enterprise AI portfolio at qwen.ai. Lasting returns come from treating AI like infrastructure: planned over years, delivered in stages, with Qwen alongside you so the early choices hold up and the value keeps growing.
About Qwen. Qwen gives enterprises an open route into production AI: one model family covering language, vision, audio, coding, embeddings, and agents, as open weights and managed APIs. Built on curated data and ongoing research, Qwen supports creating, testing, and running AI systems for decisions that matter, from self-serve developer access to dedicated inference.
Sources: McKinsey, The State of AI in 2026: On the Road to ROI (Aug 2026); Alibaba Cloud Model Studio, Model pricing (accessed Oct 5, 2026); Alibaba Cloud Model Studio, Text embedding (accessed Oct 5, 2026); Alibaba Cloud Model Studio, Knowledge base (accessed Oct 5, 2026); Alibaba Cloud Model Studio, Model training overview (accessed Oct 5, 2026); Alibaba Cloud Model Studio, Model evaluation (accessed Oct 5, 2026); Alibaba Cloud Model Studio, Privacy notice (accessed Oct 5, 2026).
Johnny Mai - Director, Gen AI Product Strategy & GTM, Qwen.
Disclaimer: This post is provided for general information only. Worked examples are illustrative and use stated assumptions; your costs will vary. Prices are Alibaba Cloud Model Studio list prices for the Singapore region as published in October 2026, vary by region, and may change over time. Figures from third-party research (McKinsey 2026) are quoted as published in those reports and are not independently verified by Qwen or Alibaba Cloud. Product and service availability, features, and program terms vary by region and may change over time. Nothing in this post constitutes professional, legal, or investment advice. © 2026 Qwen. All rights reserved.
Should This Workflow Be an Agent? A Three-Question Test Before You Build
How to Prioritize AI Use Cases Before 2027: Six Types and One 2x2
14 posts | 2 followers
FollowFarruh - June 9, 2026
Data Geek - November 4, 2024
Neel_Shah - July 2, 2026
Alibaba Cloud Indonesia - April 14, 2025
Community Builder - August 18, 2026
Community Builder - July 1, 2026
14 posts | 2 followers
Follow
Qwen
Full-range, open-source, multimodal, and multi-functional
Learn More
Alibaba Cloud Model Studio
A one-stop generative AI platform to build intelligent applications that understand your business, based on Qwen model series such as Qwen-Max and other popular models
Learn More
Token Plan
Build more, spend less. One plan, every modality.
Learn More
Alibaba Cloud for Generative AI
Accelerate innovation with generative AI to create new business success
Learn MoreMore Posts by Johnny Mai