Model Studio is Alibaba Cloud's large-model service platform, serving millions of users who simultaneously invoke dozens of large models such as Qwen. As one of the largest public-cloud model-inference entry points in China, The Model Studio gateway processes massive volumes of heterogeneous inference requests every day—every throttling misjudgment means a timeout, a retry, or even business interruption on the user side; every overload let through means a waste of scarce GPU compute and a degraded experience for other tenants.
As platform users grow from thousands to millions, the throttling system no longer faces the old problem of “anti-abuse,” but rather “how to achieve fine-grained traffic governance for millions of tenants—each independent and elastic on demand—within a limited GPU pool.”The Model Studio gateway team answered this question with an end-to-end refactor, whose core technology choice was RocketMQ LiteTopic. After the solution went live, the throttling ratio dropped by 10×, and user-perceivable anomalies caused by throttling were substantially reduced.
The following is a complete technical retrospective of this refactor.
Over the past decade, CPU, memory, and bandwidth could all scale up within minutes or even seconds, so throttling for internet applications mainly solved “anti-abuse” and “anti-avalanche,” which was fairly straightforward.
Entering the large-model era, GPUs have long scale-out cycles and high unit prices, and their supply is still constrained by the pace of hardware delivery.For public-cloud model services, “how big the GPU pool is” almost directly determines “how many customers can be served today and how high an SLA can be promised.”
As the traffic entry point for model services, the Model Studio gateway must simultaneously deliver three things within a limited GPU pool:
To do all three well at once, the traditional paradigm of “a single rate limiter sitting on the gateway” is no longer sufficient. Over the past few months, the Model Studio gateway team carried out—spanning from the throttling algorithm down to the underlying message queue—an end-to-end refactor, taking RocketMQ LiteTopic as the physical carrier of the leaky bucket, to build a fine-grained throttling system for large-model scenarios. This article shares the design trade-offs and engineering experience behind it.
What Model Studio faces is not a single throttling scenario, but three overlapping ones:
The interweaving of the three means throttling is not just about “deciding whether to let a request through,” but also about answering “at what pace to feed the backend after letting it through.” For the algorithm, the Model Studio gateway ultimately settled on fixed window + leaky bucket:
After the approach was chosen, a new problem followed: if the burst traffic of very large customers piles up in the gateway process's memory, hundreds of thousands of requests would immediately bring down the gateway itself.The leaky bucket must be moved from in-process to out-of-process, using an external store that is large enough and well isolated to absorb the buffer, physically separating “catching the requests” from “releasing them at a controlled pace.” This is where LiteTopic came into view.
Looking only at the step of “using MQ as a leaky bucket,” traditional RocketMQ Topics can do it too—and that is exactly how Model Studio first implemented it: creating a separate Topic and Consumer Group for each User + Model, then deploying a set of consumer machines dedicated to that link. But as top-tier customers came onboard one after another, three pain points quickly surfaced.
The fact is: the traditional approach can serve “a handful of top-tier bursty customers,” but cannot support “any medium-to-large customer being automatically onboarded to the leaky bucket.”When throttling needs to change from “a privilege for a few VIPs” into “a platform baseline capability,” the underlying architecture must be redesigned.
LiteTopic is a lightweight queue form introduced in RocketMQ 5.x, and its three most critical differences from traditional Topics correspond exactly to the three pain points above.
Suspend N (in milliseconds) in the consumption callback, the business can make the Broker pause pulling from that LiteTopic for N milliseconds, without affecting the pull cadence of other LiteTopics. This is the key switch that lets the leaky bucket be realized at the Broker layer.Based on these three points, the Model Studio gateway's throttling architecture was refactored as follows:
Suspend N, otherwise forward to the model and ACK. The Broker stops delivering that LiteTopic for the specified number of milliseconds and automatically resumes when the time is up.This mechanism can be understood as a “distributed leaky-bucket matrix”: each User + Model has an independent leaky bucket, whose capacity is determined by the LiteTopic's backlog limit and TTL, and whose drain rate is dynamically determined by the Suspend duration on the consumption side. The entire matrix shares the same set of consumer machines, and the resource pool scales with total traffic rather than growing linearly with the number of customers.

Using LiteTopic as the physical carrier of the leaky bucket is not merely “swapping in another queue implementation”; it matches the “scarce GPU + multi-tenant + bursty spikes” scenario along three dimensions.
Combining the three points essentially upgrades the leaky bucket from a “passive peak-shaving tool” into a piece of infrastructure that is multi-tenant-friendly, approximately unlimited in capacity, and rate-schedulable.
The core changes can be distilled into three snippets (the following are simplified illustrations showing the core call logic):
1. Sending: set the LiteTopic when constructing the message and route automatically by customer dimension; when the LiteTopic does not exist it is created automatically by the Broker.
// Gateway side: route requests to the corresponding LiteTopic by user + model
Message msg = new Message(parentTopic, payload);
msg.setLiteTopic(buildLiteTopicName(user, model)); // newly added line
producer.send(msg);
2. Consumption: at startup, use a wildcard to subscribe to all LiteTopics under the ParentTopic; new customers and new LiteTopics are automatically perceived by the consumer group, with no manifest to maintain and no restart needed.
// Consumer side: subscribe once to cover all current and future LiteTopics
PushConsumer consumer = new PushConsumer("bailian-rate-limit-group");
consumer.subscribe("bailian-rate-limit-parent", "*"); // wildcard subscription
consumer.start();
3. Throttling logic: a few lines of code in the consumption callback—if the policy hits, Suspend(N); if not, call the model normally and ACK.
@Override
public ConsumeResult consume(MessageView msg) {
String liteTopic = msg.getLiteTopic();
long suspendMs = rateLimitPolicy.acquireOrSuspend(liteTopic);
if (suspendMs > 0) {
// Suspend fetching for this LiteTopic only; other LiteTopics remain unaffected
return ConsumeResult.Suspend(suspendMs);
}
invokeModel(msg);
return ConsumeResult.SUCCESS;
}
Having hundreds of thousands of concurrent User + Model at the same time is the daily norm. If one follows the traditional practice of “launching a separate Pull loop for each subscribed LiteTopic,” the overhead grows linearly with the number of subscriptions—essentially the same as select/poll: every call must traverse the entire set, and most of the traversal is wasted.
On the Broker side, LiteTopic introduces a “Ready Set”-based event-driven mechanism, with an approach similar to epoll: only LiteTopics into which new messages are actually written enter the ready set; on Pull, the Broker reads only the ready LiteTopics, and if the set is empty it waits via long polling for an event to be triggered (arrival of a new message, still-unconsumed messages remaining after ACK, release of an ordering lock, etc.). Pull overhead is therefore proportional only to “the number of LiteTopics that are ready at the current moment,” rather than “the total number of subscriptions.”
In POC stress testing, a single Broker carried 200 clients × 10,000 LiteTopics (2 million queues in total), with average consumption latency stable at 12ms; at the same subscription volume, traditional per-Topic polling used several times more CPU when messages were sparse, and degraded further as the number of subscriptions grew.
The current minimum granularity is 30 milliseconds; finer rate control requires the consumption side to do a local Sleep itself. This step size is sufficient for the vast majority of model-inference scenarios, since inference itself takes from hundreds of milliseconds to seconds; but if the business's target rate is very high, one needs to compute the actual release rate together with the Suspend step size and the number of concurrent threads, to avoid a staircase effect propagating to P99 RT.
Under a multi-tenant shared consumer group, all LiteTopics share the same thread pool (50 threads in the POC), so the throttling method directly determines whether other LiteTopics suffer “collateral damage.”
POC comparison: subscribe to 50 LiteTopics, with the first 5 randomly triggering 300ms throttling.
Thread.sleep(300): each throttled LiteTopic holds onto a thread and does not release it; when the number of throttled ones grows to 40, the thread pool is nearly exhausted, and the remaining 10 normal LiteTopics cannot get threads and all pile up—the throttling “infects” tenants that should not have been throttled.Suspend 300: after the consumption thread returns, it immediately releases back to the thread pool and turns to serve other LiteTopics; the suspended LiteTopic is re-delivered by the Broker after N milliseconds, and the whole process does not occupy any client thread, with the result that only the 5 throttled ones pile up while the other 45 remain normal.The essential difference: Sleep is thread-level blocking that spreads to unrelated tenants through the shared thread pool; Suspend is Broker-level pull flow control that is completely decoupled at the thread level from other leaky buckets. In the normal situation of “many users being throttled at the same time while a few users are let through normally,” Suspend's non-blocking nature guarantees fairness—no matter how many User + Model are being throttled, the remaining users always have idle threads available.
Looking back at this refactor, the Model Studio gateway went from an “in-process leaky bucket” to a “distributed leaky-bucket matrix”; the core was not how novel the algorithm was, but finding an engineering carrier that makes isolation, rate adjustment, and elasticity hold true simultaneously.
A few lessons from the process are worth referencing for teams facing the same large-model throttling challenges:
From a business-outcome standpoint, the “lightweight queue + differentiated subscription + Suspend rate control” trio provided by LiteTopic reduced the Model Studio gateway's throttling ratio by 10×.
And this solution is not custom engineering unique to Model Studio—any large-model platform or inference service will run into the same throttling dilemma as long as it faces the three constraints of “scarce GPU + massive tenants + bursty spikes.”Millions of physically isolated queues + dynamic rate scheduling + zero-ops elasticity is an infrastructure paradigm that can be directly reused.
In the large-model era, throttling is no longer a defense problem, but a resource-scheduling problem. When GPUs become the scarcest factor of production, whoever can allocate limited compute to each tenant more finely, more fairly, and more elastically will find the optimal solution between experience and cost. This is precisely the value of RocketMQ LiteTopic in this era.
Let AI Agents See the Business as It Happens — Alibaba Cloud EventHouse Now Generally Available
776 posts | 60 followers
FollowAlibaba Cloud Native Community - May 19, 2026
Alibaba Cloud Native Community - September 16, 2026
Alibaba Cloud Native Community - May 14, 2026
Alibaba Cloud Native Community - June 8, 2026
Alibaba Cloud Native Community - October 22, 2025
Alibaba Cloud Native Community - January 5, 2023
776 posts | 60 followers
Follow
ApsaraMQ for RocketMQ
ApsaraMQ for RocketMQ is a distributed message queue service that supports reliable message-based asynchronous communication among microservices, distributed systems, and serverless applications.
Learn More
ApsaraMQ for MQTT
A message service designed for IoT and mobile Internet (MI).
Learn More
Message Queue for RabbitMQ
A distributed, fully managed, and professional messaging service that features high throughput, low latency, and high scalability.
Learn More
Message Queue for Apache Kafka
A fully-managed Apache Kafka service to help you quickly build data pipelines for your big data analytics.
Learn MoreMore Posts by Alibaba Cloud Native Community