×
Community Blog Should This Workflow Be an Agent? A Three-Question Test Before You Build

Should This Workflow Be an Agent? A Three-Question Test Before You Build

Operations and product leaders are probably holding two kinds of agent requests: one to agentify a process that already works, one from a team drowning in exceptions.

Should This Workflow Be an Agent? A Three-Question Test Before You Build, Alibaba Cloud Model Studio cover

In the worked example below, the agent costs $0.1728 a task. The fixed workflow costs $0.000685. The question is not whether agents work. It is whether this workflow needs one.

The Fork, Read, Undo test: no fork and no read is a script, read without fork is a workflow with model calls, fork plus read plus undo is an agent, and no undo is never an agent

Build the cheapest rung that clears the bar, then climb only when the work forces you.

In this guide

The Fork, Read, Undo test: three questions that sort any workflow into script, workflow or agent

A four-rung ladder so you climb only when real cases force it

Cost per resolved task: the breakeven math that decides whether an agent pays

Six guardrail minimums to set before the first live call

Only 17% of organizations have deployed AI agents. More than 60% expect to within two years (Gartner, 2026 CIO and Technology Executive Survey, cited in Hype Cycle for Agentic AI, 2026). So if you run operations or own a product, you are probably holding two agent requests right now.

One is from a team that wants to "agentify" a process that already works. The other is from a team drowning in exceptions no rule set can keep up with.

Fund the first and you pay agent prices for script work. Ignore the second and the exception queue keeps growing.

The answer is a gate, not a vibe. Three questions decide whether a workflow should be an agent, a fixed workflow with model calls, or a plain script. I call them the Fork, Read, Undo test.

01 · Start from the work, not the agent: when to use AI agents

An agent is a system where the model decides the next step: which tool to call, what to read, when it is done.

A workflow is one where you decide the steps in advance and the model fills in specific ones, such as extracting fields or drafting a reply. A script has no model at all.

Script, workflow and agent differ in who decides the next step: no model, you in advance, or the model

That distinction is the whole decision. Agents earn their cost when the path through the work varies by case. When the path is the same every time, an agent is a workflow that charges you to re-decide it on every run.

So start from the work. Write down one real case end to end, then a second that went differently, then a third.

Same steps all three times: no agent problem. A different fork each time: keep reading.

Already have a workflow on the agent wishlist? Pressure-test it against three real cases with the team before anyone writes a prompt.

Describe three real cases before you describe the agent.

02 · Run the Fork, Read, Undo test

Score each question 0, 1 or 2.

1. Fork: does the next step depend on what the last step found? 0 if the steps never change. 1 if there are a few known branches you could draw as a flowchart. 2 if the branches are many, unpredictable, or depend on reading something mid-task. A flowchart you can draw is a workflow. A flowchart that keeps growing is a candidate agent.

2. Read: is the deciding input unstructured? 0 if every decision runs on structured fields. 1 if some inputs are free text but the decision rests on fields. 2 if the decision depends on reading documents, emails, screens or conversations. Rule sets that try to parse prose become brittle and expensive to maintain.

3. Undo: can a wrong step be caught and reversed before it costs you? 0 if a wrong step is irreversible or harmful (money moved, data deleted, a customer told something binding). 1 if mistakes are reversible but only with effort. 2 if mistakes are cheap to catch, for example drafts reviewed before sending, or reads with no writes.

How to read the score. Fork and Read say whether an agent is useful. Undo says whether it is safe to let it decide.

Fork plus Read of 3 or more, with Undo of at least 1, is an agent candidate. Undo of 0 is a hard stop: redesign the workflow so the risky action goes through a human or a fixed rule, then rescore.

Usefulness is Fork plus Read; permission is Undo.

03 · Climb the ladder one rung at a time

The test does not produce a yes or no; it produces a rung.

Four-rung agent ladder: script, workflow with 1 to 3 model calls, single agent often 5 or more calls, and multi-agent

The agent ladder: each rung adds flexibility and cost. (Qwen.)

Two habits keep teams honest.

First, earn each rung. Build the workflow version first, and move to an agent only when you can point at the cases it fails on.

Second, a single agent with well-described tools handles more than most teams expect. Split into multiple agents when one prompt carries too many unrelated tools or instructions, not because multi-agent sounds advanced.

Every rung you climb should be paid for by cases the rung below failed.

04 · Score five example workflows

Here is the test on five common operations workflows; your scores will differ, but the method holds.

Five workflows scored on Fork, Read and Undo: password reset is a script, invoice extraction a workflow, email triage and refund disputes single agents, vendor payments not an agent

Fork, Read, Undo scores for five example workflows. (Qwen.)

Invoice extraction reads unstructured documents, but the steps never change, so a workflow with one model call does the job. Refund disputes fork on every case and depend on reading chat history and policy text: real agent territory. But the Undo score of 1 means the agent investigates and recommends while a person approves the money.

Vendor payments fail the gate outright. None of the five needs multi-agent: you would reach for rung 3 only if, say, the dispute agent also had to run fraud checks, logistics lookups and policy interpretation with tool sets too large for one prompt.

The agent investigates; the irreversible action keeps a human or a rule.

05 · Do the agent cost-per-task math before you build

Agents cost more per task because they take more turns, and each turn re-reads a growing context.

Whether that is worth it depends on what the agent removes. Here is the math for the refund dispute workflow.

Assumptions (illustrative). 20,000 disputes a month. Prices are Alibaba Cloud Model Studio list prices, Singapore region, per 1M tokens (accessed Oct 5, 2026): qwen3.8-max $2.00 input / $6.00 output; qwen3.8-flash $0.15 input / $0.47 output.

Fixed workflow: one qwen3.8-flash call per dispute to classify and route it, 3,000 input and 500 output tokens. 3,000 x $0.15 / 1M = $0.00045, plus 500 x $0.47 / 1M = $0.000235, so $0.000685 per task.

Agent, all qwen3.8-max: 6 turns, averaging 12,000 input tokens per turn (context grows as it reads orders, chats and policy) and 800 output tokens per turn. Input: 6 x 12,000 = 72,000 x $2.00 / 1M = $0.144; output: 6 x 800 = 4,800 x $6.00 / 1M = $0.0288. So $0.1728 per task, about 252 times the workflow.

Agent, routed: 5 turns on qwen3.8-flash and 1 judgment turn on qwen3.8-max. Flash: 60,000 x $0.15 / 1M = $0.009, plus 4,000 x $0.47 / 1M = $0.00188; max: 12,000 x $2.00 / 1M = $0.024, plus 800 x $6.00 / 1M = $0.0048. So $0.03968 per task.

Model cost alone makes the agent look absurd. Now add the cost it is meant to remove: human handling.

Assume a fully loaded $4 per dispute a person investigates (your number will differ; use your own). Assume the workflow escalates 30% of disputes to a person and the agent, by actually reading the case, escalates 12%.

Monthly cost of refund disputes: $24,013.70 for the fixed workflow, $13,056.00 for the all-max agent and $10,393.60 for the routed agent

Cost per month by design, refund dispute example, stated assumptions. (Qwen.)

Result: on these assumptions the routed agent totals $10,393.60 a month, against $24,013.70 for the fixed workflow.

Take the breakeven into the build decision. The all-max agent costs $3,456.00 - $13.70 = $3,442.30 more per month than the workflow.

Each percentage point of escalations removed saves 20,000 x 1% x $4 = $800. So the agent must remove at least 4.3 points of escalations (3,442.30 / 800) to pay for itself.

The routed agent needs about 1 point ((793.60 - 13.70) / 800 = 0.97). If your pilot cannot show that reduction on a graded set of real disputes, the workflow wins.

I call this cost per resolved task: model cost plus the human cost of everything the system did not resolve, divided by tasks. Per-token prices tell you what a call costs. Cost per resolved task tells you whether to build.

Each escalation point removed saves $800; the all-max agent must remove 4.3 points and the routed agent about 1 point to pay for itself

Find out whether your agent pays for itself
Send your monthly task volume, current escalation rate and loaded handling cost: one minute, no login. The team will work through your cost per resolved task and the escalation points an agent must remove, priced against the Model Studio list. Get your agent-vs-workflow breakeven →

An agent pays for itself in escalations removed, not in tokens saved.

06 · Set the guardrail minimums before the first live call

Agents act, so guardrails come first.

The market is behind: only 21% of organizations report mature governance for AI agents (Deloitte, State of AI in the Enterprise, April 2026), and only 5% of companies have the full set of controls in place for agentic AI (BCG, Applied AI Index 2026, September 2026). These are the minimums I would not ship without:

1. Input and output checks. Screen prompts and responses for unsafe or off-policy content. Model Studio offers input and output checks through the Guardrails service, billed per token.

2. Tool permissions by risk. Read tools are open; write tools are scoped; irreversible tools (payments, deletions, binding commitments) require human approval. This is the Undo score, enforced in code.

3. A turn and spend ceiling. A hard cap on turns and tokens per task, so a looping agent stops and escalates instead of running up the bill.

4. Full tracing. Every model call, tool call and decision logged per task, so a reviewer can replay what happened.

5. A graded evaluation set. Fifty or more real cases with accepted outcomes, rerun on every prompt or model change. Model Studio Model Evaluation supports LLM-as-judge, rule-based and human scoring.

6. A human handoff. A defined path, with context attached, for low-confidence cases and any case the customer asks to escalate.

Six guardrail minimums: input and output checks, tool permissions by risk, a turn and spend ceiling, full tracing, a graded evaluation set and a human handoff

Guardrails are not a tax on the agent. They raise its Undo score, and with it, how much you let it decide.

An agent gets permissions in proportion to how easily its mistakes can be undone.

07 · When should you not build an agent?

The most useful output of the test is often "no".

Do not build an agent when:

1. The steps never change. Fork is 0. Build a workflow or a script: cheaper, faster, easier to audit.

2. The deciding inputs are already structured. Read is 0 or 1 and the rules are stable. Rules engines are not glamorous, but they are deterministic.

3. Mistakes cannot be undone and you cannot add a human in the loop. Undo is 0 with no redesign. Stop.

4. You cannot measure resolution. Without a graded set and an escalation baseline, you cannot compute cost per resolved task or defend the agent at budget time.

5. Nobody inside the company can maintain it. Gartner predicts that by 2028, 70% of enterprises will abandon agentic AI built by vendor forward-deployed engineering, citing soaring costs and organizations that cannot evolve the systems themselves (Gartner, press release, September 29, 2026).

The build-or-partner call itself is the one we broke down in Build vs. Buy, or Both? The Enterprise AI Decision. For agents, the test: if your team cannot change the prompt, the tools and the evaluation set without a vendor, you do not yet own the agent.

No measurable resolution and no internal owner means no agent.

08 · Start your agent decision this week

You can run the whole test in one working session:

1. Pick three candidate workflows. Include one your team is excited about and one that is drowning in exceptions.

2. Write three real cases for each. Score Fork, Read and Undo with the workflow owner in the room.

3. Assign a rung. Script, workflow, single agent, or multi-agent, using the ladder table.

4. Price the agent candidates. Use your volume, your escalation rate and your loaded handling cost to compute breakeven, as in section 05.

5. Build the graded set before the agent. Fifty real cases with accepted outcomes.

Five-step agent decision: pick three workflows, write three cases each, assign a rung, price the candidates, build the graded set

When a workflow earns an agent, build it code-first. AgentRun is Alibaba Cloud's agent infrastructure platform: a serverless runtime, sandboxes for code and browser actions, a model gateway with load balancing, fallback and caching, MCP tools, OpenTelemetry tracing and "token-level cost attribution", which is exactly what the cost-per-resolved-task number needs.

For self-serve API access to Qwen3.8-Max and Qwen3.8-Flash with free credits on most models, start on Qwen Cloud. Model prices are on the Model Studio pricing page.

Score first, price second, build third.


Key takeaways

Fork and Read decide usefulness; Undo decides permission. An Undo score of 0 is a hard stop, whatever else is true.

Price the agent by cost per resolved task. In the dispute example, the routed agent breaks even by removing about 1 point of escalations.

Set guardrails before the first live call. Only 21% of organizations report mature governance for AI agents (Deloitte, 2026).

Score your three candidate workflows with the Model Studio team

Bring three workflows; find the right rung for each
Name three candidates, including the one your team is most excited about and the one drowning in exceptions. The team will walk through Fork, Read and Undo scores with you, plus the guardrails each rung needs before it goes live. Score your three candidate workflows →

See more of the enterprise AI portfolio at qwen.ai. Lasting returns come from treating AI like infrastructure: planned over years, delivered in stages, with Qwen alongside you so the early choices hold up and the value keeps growing.


About Qwen. Qwen gives enterprises an open route into production AI: one model family covering language, vision, audio, coding, embeddings, and agents, as open weights and managed APIs. Built on curated data and ongoing research, Qwen supports creating, testing, and running AI systems for decisions that matter, from self-serve developer access to dedicated inference.

Sources: Gartner, Hype Cycle for Agentic AI (2026); Gartner, Predicts 70% of Enterprises Will Abandon Agentic AI Built by Vendor Forward-Deployed Engineering by 2028 (September 2026); Deloitte, State of AI in the Enterprise: AI agents scaling faster than guardrails (April 2026); BCG, Applied AI Index 2026 (September 2026); Alibaba Cloud Model Studio, Model Pricing and Model Evaluation documentation (accessed Oct 5, 2026); Alibaba Cloud, AgentRun documentation (accessed Oct 5, 2026); Alibaba Cloud, Guardrails service documentation (accessed Oct 5, 2026); Qwen Cloud (accessed Oct 5, 2026).

Johnny Mai - Director, Gen AI Product Strategy & GTM, Qwen.


Disclaimer: This post is provided for general information only. Worked examples are illustrative and use stated assumptions; your costs will vary. Prices are Alibaba Cloud Model Studio list prices for the Singapore region as published in October 2026, vary by region, and may change over time. Figures from third-party research (Gartner 2026; Deloitte 2026; BCG 2026) are quoted as published in those reports and are not independently verified by Qwen or Alibaba Cloud. Product and service availability, features, and program terms vary by region and may change over time. Nothing in this post constitutes professional, legal, or investment advice. © 2026 Qwen. All rights reserved.

0 0 0
Share on

Johnny Mai

14 posts | 2 followers

You may also like

Comments