After you install the Python agent for your Large Language Model (LLM) application, ARMS begins monitoring it. Use the token analysis page to understand token usage in your LLM application.
In LLM applications, a token is the basic unit for processing text. It represents the smallest semantic unit of a model's input and output. A token can be a word, a subword, or a character, depending on the model's tokenizer.
Prerequisites
The agent for your LLM application has been installed. For more information, see Connect an LLM application or inference service to ARMS.
View token analysis
-
Log in to the ARMS console. In the left-side navigation pane, choose .
-
On the Application List page, select the region, and then click the name of your application.
-
In the top navigation bar, click token analysis.

Panel
Description
token usage
The total number of tokens consumed by LLM invocations in the selected time range.
average tokens per LLM call
The average number of tokens consumed per LLM invocation.
average tokens per request
The average number of tokens consumed per user request.
tokens consumption per minute
The total number of tokens consumed by LLM invocations per minute.
average tokens per LLM call per minute
The average number of tokens consumed per LLM invocation per minute.
average tokens per request per minute
The average number of tokens consumed per user request per minute.
token usage model ranking (Top 5)
Displays the top 5 models with the highest token consumption, sorted in descending order.
token usage session ranking (Top 5)
Displays the top 5 sessions with the highest token consumption, sorted in descending order.
token usage user ranking (Top 5)
Displays the top 5 users with the highest token consumption, sorted in descending order.