×
Community Blog Model Studio Gateway in Practice: Using RocketMQ LiteTopic to Cut the Throttling Ratio by 10×

Model Studio Gateway in Practice: Using RocketMQ LiteTopic to Cut the Throttling Ratio by 10×

This article explains how Alibaba Cloud's Model Studio gateway rebuilt its throttling system with RocketMQ LiteTopic.

Model Studio is Alibaba Cloud's large-model service platform, serving millions of users who simultaneously invoke dozens of large models such as Qwen. As one of the largest public-cloud model-inference entry points in China, The Model Studio gateway processes massive volumes of heterogeneous inference requests every day—every throttling misjudgment means a timeout, a retry, or even business interruption on the user side; every overload let through means a waste of scarce GPU compute and a degraded experience for other tenants.

As platform users grow from thousands to millions, the throttling system no longer faces the old problem of “anti-abuse,” but rather “how to achieve fine-grained traffic governance for millions of tenants—each independent and elastic on demand—within a limited GPU pool.”The Model Studio gateway team answered this question with an end-to-end refactor, whose core technology choice was RocketMQ LiteTopic. After the solution went live, the throttling ratio dropped by 10×, and user-perceivable anomalies caused by throttling were substantially reduced.

The following is a complete technical retrospective of this refactor.

1. In the Large-Model Era, Throttling Evolves from Coarse to Fine-Grained

Over the past decade, CPU, memory, and bandwidth could all scale up within minutes or even seconds, so throttling for internet applications mainly solved “anti-abuse” and “anti-avalanche,” which was fairly straightforward.

Entering the large-model era, GPUs have long scale-out cycles and high unit prices, and their supply is still constrained by the pace of hardware delivery.For public-cloud model services, “how big the GPU pool is” almost directly determines “how many customers can be served today and how high an SLA can be promised.”

As the traffic entry point for model services, the Model Studio gateway must simultaneously deliver three things within a limited GPU pool:

  • Tenant-level traffic isolation: The request traffic of each tenant is isolated from that of the others, effectively preventing resource contention and mutual interference among multiple tenants, and avoiding a situation where one tenant's burst traffic affects the normal use of others.
  • Fine-grained quota throttling: Supports differentiated quota management at the tenant dimension, applying throttling and back-off only to over-quota requests, while tenants that have not reached their quota are not restricted in any way.
  • Smoothing bursts and coordinating resources: Through traffic shaping to shave peaks and fill valleys, resources are reasonably coordinated and allocated within a limited GPU pool, ensuring a stable, low-jitter service experience.

To do all three well at once, the traditional paradigm of “a single rate limiter sitting on the gateway” is no longer sufficient. Over the past few months, the Model Studio gateway team carried out—spanning from the throttling algorithm down to the underlying message queue—an end-to-end refactor, taking RocketMQ LiteTopic as the physical carrier of the leaky bucket, to build a fine-grained throttling system for large-model scenarios. This article shares the design trade-offs and engineering experience behind it.

2. Why Are Traditional Throttling Models No Longer Enough?

What Model Studio faces is not a single throttling scenario, but three overlapping ones:

  • Baseline throttling based on SLA commitments: needs to be stable, measurable, and observable.
  • Fine-grained control across users and across multiple accounts within a customer: requires cross-constraints along the User + Model dimensions.
  • Scale-out and burst absorption for very large customers: once the quota is raised very high, when a burst actually arrives the GPUs simply cannot hold up before scale-out is in place—hard limits will return 503 on a large scale and hurt the experience, while letting the traffic through will overload the GPUs and harm other tenants.

The interweaving of the three means throttling is not just about “deciding whether to let a request through,” but also about answering “at what pace to feed the backend after letting it through.” For the algorithm, the Model Studio gateway ultimately settled on fixed window + leaky bucket:

  • Fixed window rather than sliding window: a sliding window is stricter and precisely counts bursts at the boundaries; a fixed window is more tolerant of short-term fluctuations.
  • Leaky bucket rather than token bucket: the leaky bucket's “constant-rate release” better fits GPUs—a downstream that favors steady state and is sensitive to spikes—whereas the token bucket naturally allows bursts, which is exactly the pattern GPUs fear most.

After the approach was chosen, a new problem followed: if the burst traffic of very large customers piles up in the gateway process's memory, hundreds of thousands of requests would immediately bring down the gateway itself.The leaky bucket must be moved from in-process to out-of-process, using an external store that is large enough and well isolated to absorb the buffer, physically separating “catching the requests” from “releasing them at a controlled pace.” This is where LiteTopic came into view.

3. Traditional Topic/Group Cannot Sustain a Platform-Level Leaky Bucket

Looking only at the step of “using MQ as a leaky bucket,” traditional RocketMQ Topics can do it too—and that is exactly how Model Studio first implemented it: creating a separate Topic and Consumer Group for each User + Model, then deploying a set of consumer machines dedicated to that link. But as top-tier customers came onboard one after another, three pain points quickly surfaced.

  • Heavy metadata, slow to take effect: traditional Topics/Groups must be created in advance, and it takes on the order of tens of seconds for them to take effect, so short-lived anomalies occur when a new user's burst traffic arrives.
  • Low machine utilization, cost that balloons with the number of customers: the shared consumption model requires that clients under the same Consumer Group subscribe to exactly the same set of Topics; otherwise, subscription inconsistency will cause backlog or even loss. To isolate customer traffic, one must create a separate Group for each customer—or even split off separate machines; even if a customer has not a single message today, the set of machines under its name must still stay resident.
  • The collateral-damage effect of a single Topic surging: if data from multiple customers is forcibly stuffed into the same Topic, once one user's messages surge, almost all consumer threads are occupied by it and all other users' messages are blocked. The team ultimately had to fall back to “one set of machines per customer,” trading machine scale for isolation strength.

The fact is: the traditional approach can serve “a handful of top-tier bursty customers,” but cannot support “any medium-to-large customer being automatically onboarded to the leaky bucket.”When throttling needs to change from “a privilege for a few VIPs” into “a platform baseline capability,” the underlying architecture must be redesigned.

4. LiteTopic: Turning the Throttling “Heavy Asset” into a “Light Asset”

LiteTopic is a lightweight queue form introduced in RocketMQ 5.x, and its three most critical differences from traditional Topics correspond exactly to the three pain points above.

  • Lightweight metadata: a single Broker supports millions of LiteTopics, created on demand at runtime and automatically reclaimed by TTL. A client can send a message to a nonexistent LiteTopic and it is created automatically; after a period of inactivity, the Broker cleans it up automatically.
  • Differentiated subscription: under the same Consumer Group, different consumer instances can subscribe to different subsets of LiteTopics without causing backlog or loss. The subscription relationship changes from a “rigid constraint” into “on-demand routing.”
  • Suspend consumption control: by returning Suspend N (in milliseconds) in the consumption callback, the business can make the Broker pause pulling from that LiteTopic for N milliseconds, without affecting the pull cadence of other LiteTopics. This is the key switch that lets the leaky bucket be realized at the Broker layer.

Based on these three points, the Model Studio gateway's throttling architecture was refactored as follows:

  • Sending side: after passing baseline throttling checks, a request is written to the corresponding LiteTopic by User + Model, with the User and Model information encoded directly into the name—naturally providing multi-tenant physical isolation.
  • Consumption side: the dozens of machine groups originally split by customer are merged into a single unified group of Pods that subscribe to all LiteTopics via wildcard; when a new customer is added, the sending side creates it automatically and the consumption side pulls it automatically—with no code changes, no restarts, and no need to add machines.
  • Throttling logic: converges into a few lines of decision-making within the consumption thread—if it hits, return Suspend N, otherwise forward to the model and ACK. The Broker stops delivering that LiteTopic for the specified number of milliseconds and automatically resumes when the time is up.

This mechanism can be understood as a “distributed leaky-bucket matrix”: each User + Model has an independent leaky bucket, whose capacity is determined by the LiteTopic's backlog limit and TTL, and whose drain rate is dynamically determined by the Suspend duration on the consumption side. The entire matrix shares the same set of consumer machines, and the resource pool scales with total traffic rather than growing linearly with the number of customers.

1

5. Why Is LiteTopic a Natural Carrier for the Leaky Bucket?

Using LiteTopic as the physical carrier of the leaky bucket is not merely “swapping in another queue implementation”; it matches the “scarce GPU + multi-tenant + bursty spikes” scenario along three dimensions.

  • The bucket's capacity is, in engineering terms, approximately unlimited: traditional leaky buckets are implemented with in-process queues whose capacity is a hard-coded number, and the bucket rejects once full. LiteTopic persists data to the Broker's disk, a single instance carries millions of queues, its backlog limit far exceeds that of an in-process queue, and the queues are physically isolated from one another. Customer spikes are “caught, accumulated, and released at a controlled pace,” and “processing a few seconds later” is, in almost all scenarios, better than “being 429'd.”
  • Each User + Model has its own independent leaky bucket: LiteTopic turns tenant isolation into a simple “naming rule” problem—customer A's Model X goes into liteTopic-A-X, customer B's Model X goes into liteTopic-B-X, and buckets are physically isolated at the Broker layer, so a blockage on A's link will not spread to B through any shared resource. Tenant isolation changes from “maintained by business code” into “naturally provided by infrastructure.”
  • Each leaky bucket's rate can be adjusted independently and dynamically: this is the most critical capability. Suspend N is computed entirely in real time by business policy, allowing Model Studio to adjust the release speed of any leaky bucket at the millisecond level: important customers can still obtain a higher cadence when model load is high; once model scale-out is in place, the matrix rate is raised in sync; if a certain model shows early signs of congestion, one can tighten just that column of leaky buckets alone, without dragging down other models. The whole process requires no consumer-group restart, no changes to LiteTopic configuration, and no customer awareness.

Combining the three points essentially upgrades the leaky bucket from a “passive peak-shaving tool” into a piece of infrastructure that is multi-tenant-friendly, approximately unlimited in capacity, and rate-schedulable.

6. Minimal Changes: One Snippet Each for Sending, Subscribing, and Throttling

The core changes can be distilled into three snippets (the following are simplified illustrations showing the core call logic):

1. Sending: set the LiteTopic when constructing the message and route automatically by customer dimension; when the LiteTopic does not exist it is created automatically by the Broker.

// Gateway side: route requests to the corresponding LiteTopic by user + model
Message msg = new Message(parentTopic, payload);
msg.setLiteTopic(buildLiteTopicName(user, model)); // newly added line
producer.send(msg);

2. Consumption: at startup, use a wildcard to subscribe to all LiteTopics under the ParentTopic; new customers and new LiteTopics are automatically perceived by the consumer group, with no manifest to maintain and no restart needed.

// Consumer side: subscribe once to cover all current and future LiteTopics
PushConsumer consumer = new PushConsumer("bailian-rate-limit-group");
consumer.subscribe("bailian-rate-limit-parent", "*"); // wildcard subscription
consumer.start();

3. Throttling logic: a few lines of code in the consumption callback—if the policy hits, Suspend(N); if not, call the model normally and ACK.

@Override
public ConsumeResult consume(MessageView msg) {
    String liteTopic = msg.getLiteTopic();
    long suspendMs = rateLimitPolicy.acquireOrSuspend(liteTopic);
    if (suspendMs > 0) {
        // Suspend fetching for this LiteTopic only; other LiteTopics remain unaffected
        return ConsumeResult.Suspend(suspendMs);
    }
    invokeModel(msg);
    return ConsumeResult.SUCCESS;
}

7. A Few Key Engineering Details

1. Consumption Performance Under Massive Numbers of LiteTopics

Having hundreds of thousands of concurrent User + Model at the same time is the daily norm. If one follows the traditional practice of “launching a separate Pull loop for each subscribed LiteTopic,” the overhead grows linearly with the number of subscriptions—essentially the same as select/poll: every call must traverse the entire set, and most of the traversal is wasted.

On the Broker side, LiteTopic introduces a “Ready Set”-based event-driven mechanism, with an approach similar to epoll: only LiteTopics into which new messages are actually written enter the ready set; on Pull, the Broker reads only the ready LiteTopics, and if the set is empty it waits via long polling for an event to be triggered (arrival of a new message, still-unconsumed messages remaining after ACK, release of an ordering lock, etc.). Pull overhead is therefore proportional only to “the number of LiteTopics that are ready at the current moment,” rather than “the total number of subscriptions.”

In POC stress testing, a single Broker carried 200 clients × 10,000 LiteTopics (2 million queues in total), with average consumption latency stable at 12ms; at the same subscription volume, traditional per-Topic polling used several times more CPU when messages were sparse, and degraded further as the number of subscriptions grew.

2. The Precision of Suspend

The current minimum granularity is 30 milliseconds; finer rate control requires the consumption side to do a local Sleep itself. This step size is sufficient for the vast majority of model-inference scenarios, since inference itself takes from hundreds of milliseconds to seconds; but if the business's target rate is very high, one needs to compute the actual release rate together with the Suspend step size and the number of concurrent threads, to avoid a staircase effect propagating to P99 RT.

3. Suspend vs. Consumer-Thread Sleep

Under a multi-tenant shared consumer group, all LiteTopics share the same thread pool (50 threads in the POC), so the throttling method directly determines whether other LiteTopics suffer “collateral damage.”

POC comparison: subscribe to 50 LiteTopics, with the first 5 randomly triggering 300ms throttling.

  • Thread.sleep(300): each throttled LiteTopic holds onto a thread and does not release it; when the number of throttled ones grows to 40, the thread pool is nearly exhausted, and the remaining 10 normal LiteTopics cannot get threads and all pile up—the throttling “infects” tenants that should not have been throttled.
  • Suspend 300: after the consumption thread returns, it immediately releases back to the thread pool and turns to serve other LiteTopics; the suspended LiteTopic is re-delivered by the Broker after N milliseconds, and the whole process does not occupy any client thread, with the result that only the 5 throttled ones pile up while the other 45 remain normal.

The essential difference: Sleep is thread-level blocking that spreads to unrelated tenants through the shared thread pool; Suspend is Broker-level pull flow control that is completely decoupled at the thread level from other leaky buckets. In the normal situation of “many users being throttled at the same time while a few users are let through normally,” Suspend's non-blocking nature guarantees fairness—no matter how many User + Model are being throttled, the remaining users always have idle threads available.

8. Final Thoughts

Looking back at this refactor, the Model Studio gateway went from an “in-process leaky bucket” to a “distributed leaky-bucket matrix”; the core was not how novel the algorithm was, but finding an engineering carrier that makes isolation, rate adjustment, and elasticity hold true simultaneously.

A few lessons from the process are worth referencing for teams facing the same large-model throttling challenges:

  • Throttling is not a single algorithm, but a layered combination. The fixed window manages the hard upper limit, the leaky bucket manages the release cadence, and the message queue manages the absorption buffer—the three layers are chained together and none can be missing.
  • The physical carrier of the leaky bucket determines the upper limit. When the throttling quota grows large enough that burst traffic could crush the gateway process, the leaky bucket must be moved from in-process to out-of-process. MQ is a natural candidate, but one must probe further: can it support millions of queues, can it create them on demand at runtime, and is the flow-control granularity fine enough.
  • LLM scenarios demand higher tenant isolation, but the cost of isolation must be lower. A single large-model inference request consumes several seconds of GPU time; being squeezed out loses not only success rate but also the expensive compute already consumed, so isolation becomes a hard requirement. But if it relies on “carving out a separate set of machines for each tenant,” cost balloons multiplicatively with the number of customers, which the public cloud cannot sustain. LiteTopic's solution is: each User + Model has a physically isolated queue, all queues share the same set of consumer Pods, the resource pool scales with total traffic rather than ballooning with the number of customers, and isolation strength does not drop while cost changes from multiplication back to addition.
  • Suspend and Sleep each have their applicable scenarios; the key is “who pays for the wait.” Suspend suits coarse-grained, long-duration throttling waits—pausing the entire LiteTopic's pulling, releasing the thread immediately, and not affecting other tenants; Sleep suits short-duration processing waits—for example, doing sub-millisecond cadence control when Suspend's 30ms precision is not enough, with an extremely short duration and controllable impact on the thread pool. In practice, Model Studio combines the two: Suspend takes on the main leaky-bucket rate adjustment, while Sleep supplements in the very few sub-30ms-precision scenarios.

From a business-outcome standpoint, the “lightweight queue + differentiated subscription + Suspend rate control” trio provided by LiteTopic reduced the Model Studio gateway's throttling ratio by 10×.

And this solution is not custom engineering unique to Model Studio—any large-model platform or inference service will run into the same throttling dilemma as long as it faces the three constraints of “scarce GPU + massive tenants + bursty spikes.”Millions of physically isolated queues + dynamic rate scheduling + zero-ops elasticity is an infrastructure paradigm that can be directly reused.

In the large-model era, throttling is no longer a defense problem, but a resource-scheduling problem. When GPUs become the scarcest factor of production, whoever can allocate limited compute to each tenant more finely, more fairly, and more elastically will find the optimal solution between experience and cost. This is precisely the value of RocketMQ LiteTopic in this era.

0 1 0
Share on

You may also like

Comments

Related Products