If an Elastic High Performance Computing (E-HPC) cluster or its nodes is in the Exception state, you can restore the cluster. This topic describes how to restore a cluster.
Prerequisites
-
The cluster restoration feature is enabled. By default, the feature is disabled. To enable it, submit a ticket.
-
Job data is exported.
Precautions
When you restore a cluster, take note of the following impacts:
-
When a cluster is being restored, the system disks of all nodes are changed. By default, new system disks are configured based on the settings that you specified when the cluster was created.
-
The self-managed queues in the cluster are deleted. All nodes are retained and migrated to the default queue of the cluster.
-
After a cluster is restored, the data on the system disks and data disks of all cluster nodes is lost. The data includes user information, job information, scheduler queue information, and configuration data of auto-scaling queues. However, the data on File Storage NAS file systems is retained.
Procedure
-
Log on to the E-HPC console.
-
In the left part of the top navigation bar, select a region.
-
In the navigation pane on the left, click Cluster.
-
On the Cluster page, find the target cluster and choose More > Recover.
-
In the dialog box, configure the ImageType, Image, Scheduler, and Domain Service for the cluster.
Other settings are configured based on the settings that you specified when you created the cluster.
-
Click OK.
Result
After the recovery begins, the console redirects you to the cluster list page. During recovery, the cluster's status is Uninitialized. The status changes to Running when the recovery is complete.