JindoTable provides the archiveTable and unarchiveTable commands in SDK mode. You can use these commands to archive and restore data without relying on the Jindo Namespace Service. This topic describes how to use the archiveTable and unarchiveTable commands.
Prerequisites
- Java Development Kit (JDK) 8 is installed locally.
- A cluster is created. For more information, see Create a cluster.
- The data to be archived must be in a table, which can be a partitioned or non-partitioned table, and must be stored in Alibaba Cloud Object Storage Service (OSS).
Background information
The original JindoTable archive and unarchive commands can archive or restore tables or partitions on OSS. However, these commands depend on the Jindo Namespace Service of the SmartData component. The new archiveTable and unarchiveTable commands allow you to perform archive and restore operations without this dependency.
- You can run the commands on a cluster where the SmartData service is not deployed, such as a self-managed cluster that is not an EMR cluster.
- You can pass filter parameters to apply the operation to multiple partitions and execute it in a multi-threaded manner. If local multi-threading is insufficient, you can also start a MapReduce job to run on the entire cluster.
For more information about the original archive and unarchive commands, see JindoTable user guide.
Limits
The new archiveTable and unarchiveTable commands are supported on clusters of EMR V3.36.0 and later, or EMR V5.2.0 and later.
The archiveTable command
The archiveTable command archives tables or partitions on OSS.
- Log on to the cluster using the Secure Shell (SSH) protocol. For more information, see Log on to a cluster.
- Run the following command to obtain help information.
jindo table -help archiveTableThe following information is returned.<dbName.tableName> The table to archive. -a/-i storage policy, -a for Archive and -i for IA (Infrequent Access). <condition>/-fullTable A filter condition to determine which partitions should be archived, supporting common operators (like '>'), while -fullTable means that all partitions (or a whole un-partitioned table) should be archived. One but only one option must be specified among -c "<condition>" and -fullTable. <before days> Optional, saying that table/partitions should be archived only when they are created (not updated or modified) more than some days before from now. <parallelism> The maximum concurrency when archiving partitions, 1 by default. -mr/-mapReduce Archive table/partitions using cluster-level MapReduce job instead of local-level multi-thread. -e/-explain If present, the command would not really archive data, but only prints the table/partitions that would be archived for given conditions. <working directory>: A directory to locate map-reduce temp files. Must not be a local file system directory. 'hdfs:///tmp/<current user>/jindotable-policy/' by default. <log directory> A directory to locate log files, '/tmp/<current user>/' by default.The syntax of the archiveTable command is as follows.-archiveTable -t <dbName.tableName> \ -a/-i \ [-c "<condition>" | -fullTable] \ [-b/-before <before days>] \ [-p/-parallel <parallelism>] \ [-mr/-mapReduce] \ [-e/-explain] \ [-w/-workingDir <working directory>] \ [-l/-logDir <log directory>]Parameter Description Required -t <dbName.tableName> The name of the table to archive. The format is database_name.table_name.Use a period (.) to separate the database name and the table name. The table can be a partitioned or non-partitioned table.
Yes -a/-i The destination storage class. The following options are supported: -a: Archive Storage.-i: Infrequent Access (IA) storage.
If you specify -i for IA storage, files that are already in Archive Storage are skipped.
Yes -c "<condition>" | -fullTable -fullTableor-c "<condition>".- If you specify
-fullTable, the entire table is moved. The table can be partitioned or non-partitioned. - If you specify
-c "<condition>", a filter condition is provided to select the partitions to move. Common operators, such as the greater-than sign (>), are supported.For example, for a partition key column `ds` of the String data type, if you want to select partitions where the partition name is greater than 'd', use
-c " ds > 'd' ".
No -b/before <before days> Only tables or partitions created more than the specified number of days ago are archived. No -p/-parallel <parallelism> The degree of parallelism for the archive operation. No -mr/-mapReduce Uses Hadoop MapReduce instead of local multi-threading to archive data. No -e/-explain If this option is present, the command runs in explain mode. It only displays the list of partitions to be moved and does not actually move any data. No -w/-workingDir Used only for MapReduce jobs. This specifies the working directory for the MapReduce job. You must have read and write permissions on the directory. The working directory can be non-empty. Temporary files are created during job execution and are cleaned up after the job is complete. No -l/-logDir <log directory> Specifies the directory for log files. No
The unarchiveTable command
The unarchiveTable command has a format similar to the archiveTable command and performs the opposite operation. It restores tables or partitions on OSS.
- Log on to the cluster using the SSH protocol. For more information, see Log on to a cluster.
- Run the following command to obtain help information.
jindo table -help unarchiveTableThe following information is returned.<dbName.tableName> The table to unarchive. -i unarchive to IA (Infrequent Access). -o restore to make archived data accessible temporarily. <condition>/-fullTable A filter condition to determine which partitions should be unarchived, supporting common operators (like '>'), while -fullTable means that all partitions (or a whole un-partitioned table) should be unarchived. One but only one option must be specified among -c "<condition>" and -fullTable. <before days> Optional, saying that table/partitions should be unarchived only when they are created (not updated or modified) more than some days before from now. <parallelism> The maximum concurrency when unarchiving partitions, 1 by default. -mr/-mapReduce Unarchive table/partitions using cluster-level MapReduce job instead of local-level multi-thread. -e/-explain If present, the command would not really unarchive data, but only prints the table/partitions that would be unarchived for given conditions. <working directory>: A directory to locate map-reduce temp files. Must not be a local file system directory. 'hdfs:///tmp/<current user>/jindotable-policy/' by default. <log directory> A directory to locate log files, '/tmp/<current user>/' by default.The syntax of the unarchiveTable command is as follows.-unarchiveTable -t <dbName.tableName> \ [-i/-o] \ [-c "<condition>" | -fullTable] \ [-b/-before <before days>] \ [-p/-parallel <parallelism>] \ [-mr/-mapReduce] \ [-e/-explain] \ [-w/-workingDir <working directory>] \ [-l/-logDir <log directory>]
The parameters for the unarchiveTable command are almost the same as those for the archiveTable command. The main difference is that the required parameter -a/-i is replaced by the optional parameter -i/-o.
- If you do not specify the -i/-o parameter, the storage class is converted to Standard.
- If you specify the -i parameter, the storage class is converted to Infrequent Access (IA). Files already in the Standard storage class are skipped.
- If you specify the -o parameter, only a restore operation is performed. Files that are already in the Standard or IA storage class are skipped. Files that are already in a restored state are also skipped to prevent repeated restores.