After you install the Python agent for your large language model (LLM) application, ARMS begins monitoring the application. On the performance analysis page, you can view information such as the number of model calls, average model call time, and the number of model call errors.
Prerequisites
You must have an agent installed for the LLM application. For more information, see Connect an LLM application or inference service to ARMS.
LLM application performance analysis
-
Log on to the ARMS console. In the left-side navigation pane, choose .
-
On the Application List page, select a region at the top of the page and click the name of the target application.
-
In the top navigation bar, click Performance analysis.
Panel
Description
Number of model calls
The number of calls to the large language model within the specified time period.
Average model call time
The average duration of calls to the large language model within the specified time period.
Number of model call errors
The number of failed calls to the large language model within the specified time period.
Number of model calls/min
The number of calls to the large language model per minute.
Average model call time/min
The average duration of calls to the large language model per minute.
Model call errors/min
The number of failed calls to the large language model per minute.
Model latency P99/min
The P99 latency for calls to the large language model per minute. This means that 99% of calls are completed in less than this value.
Average time to first packet/min
The average time to receive the first data packet from the large language model per minute.
P99 time to first packet/min
The P99 time to receive the first data packet from the large language model per minute.
Top 5 models by calls
The top 5 models with the most calls, in descending order.
Top 5 models by average call time
The top 5 models with the longest average call time, in descending order.
Top 5 models by call errors
The top 5 models with the most call errors, in descending order.