All Products
Search
Document Center

Elasticsearch:Event Hub

Last Updated:Jun 17, 2026

Use the Event Center to view system O&M events for Alibaba Cloud Elasticsearch (ES), quickly detect service anomalies, and pinpoint issues.

Event categories

ES events fall into the following categories based on their cause and impact.

Note

For more information, see Appendix: Event details.

Event category

Definition

Cause and impact

Examples

System change

Alibaba Cloud initiates system change events and notifies you. Check whether your cluster is affected.

Infrastructure changes or faults may affect cluster access. When such an event occurs, the system sends a notification. Check the notification and your cluster status promptly.

  • A Kibana feature upgrade causes a brief service suspension.

  • Alibaba Cloud upgrades an AMD instance family to the latest generation.

Cluster health

The system periodically inspects and monitors cluster health based on actual usage, and reports unexpected diagnostic results as events.

To ensure service continuity, the system automatically triggers a cluster health event when it detects an anomaly or risk in cluster resources.

Note

During the execution of an O&M event, the cluster may experience brief jitter without affecting normal access. If automatic execution fails, you can manually trigger a node restart on the Event Center page. You have 24 to 48 hours to intervene manually. For specific execution times, see View and handle events.

An inspection finds that an ES node is offline.

Cluster change

These events correspond to cluster changes that you initiate. Failures or blockages may occur during the change process.

Instance type changes or kernel upgrades trigger a restart of the corresponding nodes. During the restart, the cluster may experience brief jitter without affecting normal access.

  • Scale-in

  • Restart a node

View and handle events

On the Event Center page, you can view and respond to events for your account.

  1. Go to the Event Center.

    1. Log on to the Alibaba Cloud Elasticsearch console.

    2. In the navigation pane, click Event Center.

  2. View event information.

    On the Event Center page, you can filter events by type to view all events for a specific instance within a specified time period, and then respond based on the event details. The page contains three tabs: System Change, Cluster Health, and Cluster Changes. At the top of the page, use the time range selector or search by instance ID to filter events. In the upper-right corner, click Event Subscription or Manage Notifications. In the event list, click Restart or Schedule Restart in the Suggestion column to handle pending events.

    Note

    You can view all event information in the Event Center. You can also subscribe to events and set up notifications for critical alerts. When an alert is triggered, the system sends a notification to the specified contacts by phone call, text message, or email.

    The following table describes the event information and related actions.

    Event information

    Description

    Cluster ID

    The ID of the Alibaba Cloud ES instance where the event occurred.

    Node ID

    The ID of the node within the instance where the event occurred.

    Event Level

    The severity of the event. Valid values:

    • Info: Records routine system operations and status. Useful for monitoring or debugging.

    • Warning: Indicates a potential issue that does not currently affect operations but requires monitoring.

    • Critical: A serious error or fault has occurred. Immediate action is required to prevent service disruption or data loss.

    Event Status

    The execution status of the event. Valid values include To Be Handled, In Progress, Handled, Handling Failed, Handling Interrupted, Canceled, Execution to be confirmed, Ready to continue, Occurred, In Progress, and Recovered. The following describes key statuses:

    • To Be Handled: The event is waiting to be executed at the system-set time or at a time you have scheduled.

    • Execution to be confirmed: Based on the event details, you can decide whether to execute the event immediately or create a snapshot backup.

      Note
      • This status is supported only for some events related to local disks on the System Change tab.

      • Snapshot backups are available only for deployment events, such as an Alibaba Cloud ES cluster upgrade or a new version deployment to a specific node.

    • Ready to continue: The grayscale change is complete. You must confirm the stability of the affected nodes and cluster before continuing. For example, after a change is tested and verified on a few nodes, it is then applied to all remaining nodes.

    For events with a status of Handling Failed or Handling Interrupted, identify the cause and resolve the issue promptly to avoid impacting your business operations.

    Event Description

    The cause and impact of the event.

    Occurred At and Ended At

    The start and end times of the event.

    Scheduled Handling Time and Execution End Time

    The scheduled start time and estimated end time for the event handling.

    Note

    This information is available only for system change events.

    Source

    The source of the event. Valid values:

    • Proactive Notification: Alibaba Cloud ES automatically sends generated events to the Event Center.

    • Event Subscription: You subscribe to specific events. When a subscribed event occurs, you receive a notification.

    Suggestion

    Handle events based on the provided suggestions. Supported actions vary by event. Refer to the UI for details.

    • Contact Technical Support: Contact Technical Support if you have questions about an event.

    • Restart: Immediately restarts the specified node.

    • Schedule Restart: Specify a restart time. The scheduled time must be at least 5 minutes in the future. The system restarts the specified node within 5 minutes after the scheduled time.

    Note

    When you perform a restart, forced restart, or grayscale restart on an instance or node, the system triggers a corresponding restart event. For redeployment events, such as an Alibaba Cloud ES version upgrade, submit a ticket to Technical Support.

Appendix: Event details

Event type

Event code and name

CloudMonitor event name

Cause category

Event level

Description and impact

System change event

  • SystemUpdate.InfraDiskError

  • System change event due to an infrastructure disk failure

  • Instance:SystemUpdate.InfraDiskError:Executing: System change event in progress due to an infrastructure disk failure

  • Instance:SystemUpdate.InfraDiskError:Executed: System change event completed due to an infrastructure disk failure

Critical

An infrastructure failure makes the local disk unavailable.

This event requires a backend redeployment. To resolve this, submit a ticket to technical support.

  • SystemUpdate.InfraDiskStalled

  • System change event due to infrastructure disk performance issues

  • Instance:SystemUpdate.InfraDiskstalled:Executing: System change event in progress due to infrastructure disk performance issues

  • Instance:SystemUpdate.InfraDiskstalled:Executed: System change event completed due to infrastructure disk performance issues

Critical

An infrastructure failure degrades cloud disk performance.

  • SystemUpdate.InfraFailureStop

  • System change event due to an instance stop caused by an infrastructure failure

  • Instance:SystemUpdate.InfraFailureStop:Scheduled: System change event scheduled to stop the instance due to an infrastructure failure

  • Instance:SystemUpdate.InfraFailureStop:Executing: System change event in progress to stop the instance due to an infrastructure failure

  • Instance:SystemUpdate.InfraFailureStop:Executed: System change event completed to stop the instance due to an infrastructure failure

  • Instance:SystemUpdate.InfraFailureStop:Failed: System change event failed to stop the instance due to an infrastructure failure

Critical

The instance may stop due to a potential infrastructure failure.

  • SystemUpdate.InfraMigrate

  • System change event due to an infrastructure migration or upgrade

  • Instance:SystemUpdate.InfraMigrate:Scheduled: System change event scheduled for an infrastructure migration or upgrade

  • Instance:SystemUpdate.InfraMigrate:Executing: System change event in progress for an infrastructure migration or upgrade

  • Instance:SystemUpdate.InfraMigrate:Executed: System change event completed for an infrastructure migration or upgrade

  • Instance:SystemUpdate.InfraMigrate:Failed: System change event failed for an infrastructure migration or upgrade

Critical

  • The instance node restarts due to infrastructure maintenance.

  • The instance node is redeployed due to infrastructure maintenance.

  • SystemUpdate.SoftwareRepair

  • System change event due to a control system software update

  • Instance:SystemUpdate.SoftwareRepair:Scheduled: System change event scheduled for a software update

  • Instance:SystemUpdate.SoftwareRepair:Executing: System change event in progress for a software update

  • Instance:SystemUpdate.SoftwareRepair:Executed: System change event completed for a software update

Warning

  • Description: The cluster control system restarts due to an upgrade. This upgrade involves changes to the Alibaba Cloud instance architecture, which upgrades the control deployment mode from Basic Control (v2) to Cloud-native Control (v3).

    Note

    You can view the control deployment mode on the Basic Information page of the instance.

  • Impact:

    • The upgrade uses a blue-green deployment within a scheduled period. During this process, the number of cluster nodes doubles, but no extra fees are charged.

    • The upgrade takes several hours, depending on the data volume. The system takes the old nodes offline during your configured O&M window, causing a service interruption of about 1 to 2 seconds. Instance change operations are unavailable during the upgrade. Prepare your services in advance.

    • Clusters are upgraded from version 6.8.6 to 6.8.23. The engine is fully compatible, and your services are not affected.

    • After the upgrade, the Kibana private network is disabled. You need to log on to the Kibana console to enable it.

Cluster health event

  • HealthCheck.ClusterAbnormal

  • Cluster health event due to an abnormal cluster status

  • Instance:HealthCheck.ClusterAbnormal:Executed: Cluster health event completed due to an abnormal cluster status

  • Instance:HealthCheck.ClusterAbnormal:Failed: Cluster health event failed due to an abnormal cluster status

Critical

The instance restarts due to an abnormal cluster status.

  • HealthCheck.ClusterUnhealthy

  • Cluster health event due to an unhealthy cluster status

  • Instance:HealthCheck:ClusterUnhealthy:Occurred: A health check event for an unhealthy cluster has occurred.

  • Instance:HealthCheck:ClusterUnhealthy:Persistent: A health check event for an unhealthy cluster is ongoing.

  • Instance:HealthCheck:ClusterUnhealthy:Recovered: A health check event for an unhealthy cluster has been resolved.

Cluster.StatusRed: The cluster health status changes to Red.

Critical

The cluster status is Red, indicating unassigned primary shards. Data is unavailable.

Cluster.StatusYellow: The cluster health status changes to Yellow.

Warning

The cluster status is Yellow, indicating unassigned replica shards. This reduces data redundancy.

Node.Disconnected: A cluster node is offline or disconnected.

Critical

A node is offline or disconnected, which may lead to data unavailability or performance degradation.

  • HealthCheck.JVMMemoryPressure

  • Resource anomaly event due to JVM memory pressure

  • Instance:HealthCheck:JVMMemoryPressure:Occurred

  • Instance:HealthCheck:JVMMemoryPressure:Persistent

  • Instance:HealthCheck:JVMMemoryPressure:Recovered

JVMMemory.HeapMemoryHigh: High heap memory usage

Warning

High heap memory usage may trigger a full GC.

JVMMemory.HeapMemoryCritical: Critically high heap memory usage

Critical

Heap memory is near its limit and is highly likely to cause an OutOfMemory (OOM) error.

JVMMemory.GCRateTooHigh: Frequent Old GC

Warning

Frequent Old GC affects performance.

  • HealthCheck.CPULoadHigh

  • Resource anomaly event due to high CPU load

  • Instance:HealthCheck:CPULoadHigh:Occurred

  • Instance:HealthCheck:CPULoadHigh:Persistent

  • Instance:HealthCheck:CPULoadHigh:Recovered

CPU.PersistUsageHigh: Sustained high CPU load

Warning

Sustained high CPU load slows down system responsiveness.

CPU.PersistUsageCritical: Sustained high CPU load

Critical

Sustained high CPU load slows down system responsiveness.

  • HealthCheck.DiskUsageHigh

  • Resource anomaly event due to high disk usage

  • Instance:HealthCheck:DiskUsageHigh:Occurred

  • Instance:HealthCheck:DiskUsageHigh:Persistent

  • Instance:HealthCheck:DiskUsageHigh:Recovered

Disk.UsageHigh: Disk usage alert

Warning

Insufficient disk space prevents new shards from being created. Clear space or scale up the storage.

Disk.UsageCritical: Critical disk usage

Critical

Disk usage is approaching the automatic Elasticsearch read-only threshold (95%). This affects normal data writes and requires immediate action.

Disk.IndexReadOnly: The index enters a read-only state.

Critical

Elasticsearch automatically sets the index to read-only, typically when the disk is full. This action blocks all writes.

  • HealthCheck.DiskIOBottleneck

  • Resource anomaly event due to a disk I/O bottleneck

  • Instance:HealthCheck:DiskIOBottleneck:Occurred

  • Instance:HealthCheck:DiskIOBottleneck:Persistent

  • Instance:HealthCheck:DiskIOBottleneck:Recovered

Disk.IOUtilizationHigh: High disk I/O utilization

Critical

High disk I/O utilization increases read/write latency. Scale up the disk or switch to a higher-performance disk type to resolve this.

  • HealthCheck.ThreadPoolSaturation

  • Performance bottleneck event due to thread pool saturation

  • Instance:HealthCheck:ThreadPoolSaturation:Occurred

  • Instance:HealthCheck:ThreadPoolSaturation:Persistent

  • Instance:HealthCheck:ThreadPoolSaturation:Recovered

ThreadPool.SearchQueueHigh: The search thread pool queue is congested.

Warning

Congestion in the search thread pool queue slows down query responses.

ThreadPool.SearchRejected: Search requests are rejected.

Critical

The system rejects search requests, causing user queries to fail.

ThreadPool.WriteQueueHigh: The write thread pool queue is congested.

Warning

Congestion in the write thread pool queue slows down write responses.

ThreadPool.WriteRejected: Write requests are rejected.

Critical

The system rejects write requests, causing data writes to fail.

Cluster change event

  • UserOperator.InstanceSpecModify

  • Cluster change event due to an instance type change

  • Instance:UserOperator.InstanceSpecModify:Executing: Cluster change event in progress due to an instance type change

  • Instance:UserOperator.InstanceSpecModify:Executed: Cluster change event completed due to an instance type change

Info

  • The instance restarts due to an instance type change.

  • The instance node restarts due to an instance node change.

  • UserOperator.InstanceUpdate

  • Cluster change event due to an instance change operation

  • Instance:UserOperator.InstanceUpdate:Executing: Cluster change event in progress due to an instance change operation

  • Instance:UserOperator.InstanceUpdate:Executed: Cluster change event completed due to an instance change operation

Info

  • The instance restarts due to a configuration change.

  • The instance's plugins are updated.

  • The instance's IK dictionary is hot-updated.

  • UserOperator.InstanceCoreUpdate

  • Cluster change event due to an instance kernel upgrade

  • Instance:UserOperator.InstanceCoreUpdate:Executing: Cluster change event in progress due to an instance kernel upgrade

  • Instance:UserOperator.InstanceCoreUpdate:Executed: Cluster change event completed due to an instance kernel upgrade

Info

The instance restarts due to a kernel version update.