When you define an application-level service level objective (SLO) in Alibaba Cloud Service Mesh (ASM), a set of Prometheus rules is automatically generated. This topic explains how to import these rules into your Prometheus instance to enforce your SLOs.
Prerequisites
-
You have defined an application-level service level objective (SLO).
-
Prometheus monitoring is installed in your Container Service for Kubernetes (ACK) cluster. For more information, see Open source Prometheus monitoring and Integrate a self-managed Prometheus instance for mesh monitoring.
Step 1: Import rules into Prometheus
This topic assumes you have deployed Prometheus by using the Prometheus Operator. In this mode, Prometheus configurations are managed through custom resources. To configure recording rules and alerting rules, you can create a PrometheusRule object with the app: ack-prometheus-operator and release: ack-prometheus-operator labels.
-
Whether you need these labels depends on the
ruleSelectorsettings in your Prometheus custom resource. If theruleSelectoris empty, you do not need to add the labels. Adjust the configuration to match your environment. -
If you use a different method to deploy Prometheus, you must use the corresponding method to apply the generated rules. For more information, see the official Prometheus documentation.
-
Obtain the
ruleSelectorfrom the ACK console.-
Log on to the ACK console. In the left navigation pane, click Clusters.
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
-
On the CRDs tab, click PrometheusRule.
-
On the Resource Objects tab, select monitoring from the Namespace drop-down list. In the row for ack-prometheus-operator-prometheus, click Edit YAML in the Actions column.
-
Obtain the
ruleSelectorfield.The following example shows a typical
ruleSelectorconfiguration. For the Prometheus Operator to select yourPrometheusRuleresource, the resource must have the labels defined inmatchLabels.ruleSelector: matchLabels: app: ack-prometheus-operator release: ack-prometheus-operator
-
-
Deploy the PrometheusRule resource.
-
Create a file named prometheusrule.yaml with the following content.
In the YAML file, the
labelsfield must contain the labels that you obtained in the previous step. Thespecfield contains the generated Prometheus rules.apiVersion: monitoring.coreos.com/v1 kind: PrometheusRule metadata: labels: app: ack-prometheus-operator release: ack-prometheus-operator name: asm-rules namespace: monitoring spec: # Replace this with the content of your generated rule file. -
Run the following command in your ACK cluster to apply the PrometheusRule resource.
kubectl apply -f prometheusrule.yaml
-
-
Verify that the rules are applied.
-
Log on to the ACK console. In the left navigation pane, click Clusters.
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
-
On the ConfigMaps page, select monitoring from the Namespace drop-down list. In the row for the Prometheus configuration, click Edit YAML in the Actions column.
The
PrometheusRulecontroller automatically converts the resource into rules and adds them to a ConfigMap for Prometheus to read. The following example shows the ConfigMap content, wheremonitoring-asm-rules-fcd9bea9-22b3-462b-8786-424f74f098ec.yamlis the ASM rule file key:up{job= alertmanager main,namespace= monitoring } ) ) >= 0.5 for: 5m labels: severity: critical monitoring-asm-rules-fcd9bea9-22b3-462b-8786-424f74f098ec.yaml: | groups: - name: sloth-slo-sli-recordings-httpbin-asm-slo rules:
-
Step 2: Verify SLO monitoring
Check metrics and alerts in Prometheus
-
Connect to your ACK cluster using kubectl and run the following command to forward the Prometheus service to a local port.
kubectl --namespace monitoring port-forward svc/ack-prometheus-operator-prometheus 9090 -
Click https://localhost:9090 to access the Prometheus console.
-
In the Prometheus UI, enter asm_slo_info in the expression browser and click Execute to view your SLO configuration.
The query returns the result
asm_slo_info{asm_slo="asm-slo",slo_id="httpbin-asm-slo",slo_mode="cli-gen-prom",slo_objective="99.9",slo_service="httpbin",slo_spec="prometheus/v1",slo_version="dev"}, indicating that your Prometheus recording rules are configured correctly. -
At the top of the page, click Alerts to view the alerting rules.
The presence of the following two rules indicates that your Prometheus alerting rules are configured correctly. Two alerting rules named asm-alert are displayed, both showing 0 active (no active alerts).
Scenario 1: Simulate normal traffic
-
Run the following script to simulate traffic with a 99.5% success rate.
Replace
{ingress-gateway-ip}with the actual IP address of your ingress gateway. To obtain the IP address of the ingress gateway, see View ingress gateway information.#!/bin/bash for i in `seq 200` do if (( $i == 100 )) then curl -I http://{ingress-gateway-ip}/status/500; else curl -I http://{ingress-gateway-ip}/; fi echo "OK" sleep 0.01; done; -
Return to the Prometheus UI, enter slo:period_error_budget_remaining:ratio in the expression browser, and then click Execute. Observe the changes in your remaining error budget.
After executing the query, the graph shows the remaining error budget ratio dropping sharply from approximately 1.0 to approximately 0.93, indicating that the error budget is being consumed rapidly.
The following table describes the key SLO metrics. For more information, see Service level objective (SLO) overview.
Metric
Description
slo:period_error_budget_remaining:ratio
The remaining error budget for the 30-day SLO period.
slo:sli_error:ratio_rate30d
The average error rate over the 30-day SLO period.
slo:period_burn_rate:ratio
The burn rate over the 30-day SLO period.
slo:current_burn_rate:ratio
The current burn rate.
Scenario 2: Simulate error traffic
Manually trigger failures to test the alerting rules.
-
Run the following script to simulate traffic with a 50% success rate, which corresponds to a burn rate of 50.
Replace
{ingress-gateway-ip}with the actual IP address of your ingress gateway.#!/bin/bash for i in `seq 200` do curl -I http://{ingress-gateway-ip}/ curl -I http://{ingress-gateway-ip}/status/500; echo "OK" sleep 0.01; done; -
Return to the Alerts page in the Prometheus console to see the triggered alerts.
On the Alerts page, under the rule file path
monitoring-asm-rules.yaml, in thesloth-slo-alerts-httpbin-asm-slo2rule group, two asm-alert alerting rules are in active state (1 active):-
First rule:
sloth_severity=page,sloth_window=30m, labels includealertlabel="bbb",pagelabels="ddd",sloth_id="httpbin-asm-slo2". -
Second rule:
sloth_severity=ticket,sloth_window=2h, labels includeticketlabel="eee".
-
View alerts in the Alertmanager console
In the Prometheus framework, the Alertmanager component collects alerts generated by the Prometheus server and routes them to receivers based on your configuration.
-
Run the following command to forward the ack-prometheus-operator-alertmanager service to a local port.
kubectl --namespace monitoring port-forward svc/ack-prometheus-operator-alertmanager 9093 -
Click https://localhost:9093 to access the Alertmanager console.
-
On the Alertmanager page, click the
expand icon to view alert details.The Alertmanager alerts list shows two custom alerts of type asm-alert: one with
slo_severity=page(slo_window=30m) and another withslo_severity=ticket(slo_window=2h), indicating that the ASM SLO alerting rules have been triggered successfully.