Use JindoTable's archiveTable and unarchiveTable commands to move Hive tables and partitions between OSS storage classes — without requiring the Jindo Namespace Service component of SmartData.
Prerequisites
Before you begin, ensure that you have:
An E-MapReduce (EMR) cluster running EMR V3.36.0 or later, or EMR V5.2.0 or later
A partitioned or non-partitioned table stored in OSS (only table data can be archived)
SSH access to the cluster
Background
The original JindoTable archive and unarchive commands depend on the Jindo Namespace Service component of SmartData. The archiveTable and unarchiveTable commands remove that dependency, which means:
They work on clusters where SmartData is not deployed, including self-managed clusters.
They support filter parameters for multi-threaded archiving across large partition sets. If local multithreading is not enough, use MapReduce to distribute the work across the entire cluster.
For documentation on the original commands, see Use JindoTable.
Storage class transitions
The following table shows which storage classes each command can target and which files are skipped during the operation.
archiveTable
| Flag | Target storage class | Files skipped |
|---|---|---|
-i | Infrequent Access (IA) | Archive, Cold Archive |
-a | Archive | Cold Archive |
-ca | Cold Archive | None |
unarchiveTable
| Flag | Resulting storage class | Files skipped |
|---|---|---|
| (none) | Standard | None |
-i | Infrequent Access (IA) | Standard |
-a | Archive | Standard, IA |
-o | Unchanged (temporary restore only) | Standard, IA, previously restored files |
-cr | No change (check restore status only) | N/A |
Archive tables or partitions
Use archiveTable to move OSS table data to a lower-cost storage class.
Log in to your cluster over SSH. See Log on to a cluster.
Run the following command to see full help:
jindo table -help archiveTableArchive tables or partitions using the following syntax:
-archiveTable -t <dbName.tableName> \ -i/-a/-ca \ [-c "<condition>" | -fullTable] \ [-b/-before <before days>] \ [-p/-parallel <parallelism>] \ [-mr/-mapReduce] \ [-e/-explain] \ [-w/-workingDir <working directory>] \ [-l/-logDir <log directory>]
Parameters
| Parameter | Description | Required |
|---|---|---|
-t <dbName.tableName> | The table to archive, in database_name.table_name format. Supports partitioned and non-partitioned tables. | Yes |
-i / -a / -ca | Target storage class: -i for Infrequent Access (IA), -a for Archive, -ca for Cold Archive. See Storage class transitions for skip rules. | Yes |
-c "<condition>" / -fullTable | Scope of archiving: use -fullTable for the entire table, or -c "<condition>" to archive only matching partitions. Supports standard comparison operators. | Yes |
-b / -before <before days> | Only archive tables or partitions created at least N days ago. | No |
-p / -parallel <parallelism> | Number of parallel archiving threads. | No |
-mr / -mapReduce | Use Hadoop MapReduce instead of local multithreading. | No |
-e / -explain | Dry run: list partitions that would be archived without making any changes. | No |
-w / -workingDir <directory> | Working directory for the MapReduce job. Must have read and write permissions. Temporary files are created during the job and deleted automatically on completion. | No |
-l / -logDir <directory> | Directory for log files. | No |
Examples
Preview which partitions will be archived (dry run):
-archiveTable -t mydb.mytable -a -c " ds > 'd' " -eArchive all partitions in a table to Cold Archive:
-archiveTable -t mydb.mytable -ca -fullTableArchive partitions where the `ds` column is greater than `'d'`, created more than 30 days ago:
-archiveTable -t mydb.mytable -a -c " ds > 'd' " -b 30Unarchive tables or partitions
Use unarchiveTable to restore OSS table data from archived storage. Its syntax mirrors archiveTable, with two differences: the storage class flag (-i/-a/-o/-cr) is optional, and the -notWait flag is available for Cold Archive restores.
Log in to your cluster over SSH. See Log on to a cluster.
Run the following command to see full help:
jindo table -help unarchiveTableUnarchive tables or partitions using the following syntax:
-archiveTable -t <dbName.tableName> \ [-i/-a/-o/-cr] \ [-notWait] \ [-c "<condition>" | -fullTable] \ [-b/-before <before days>] \ [-p/-parallel <parallelism>] \ [-mr/-mapReduce] \ [-e/-explain] \ [-w/-workingDir <working directory>] \ [-l/-logDir <log directory>]
Parameters unique to unarchiveTable
All parameters from archiveTable apply. The following parameters differ or are new:
| Parameter | Description | Required |
|---|---|---|
-i / -a / -o / -cr | Target storage class after unarchiving. If omitted, data is restored to Standard. See Storage class transitions for details. | No |
-notWait | Valid only when unarchiving data. Submit the unarchive request for Cold Archive data and return immediately without waiting for completion. | No |
Examples
Preview which partitions will be restored (dry run):
-archiveTable -t mydb.mytable -a -c " ds > 'd' " -eRestore all partitions to Standard storage:
-archiveTable -t mydb.mytable -fullTableRestore partitions to Archive storage, skipping Standard and IA files:
-archiveTable -t mydb.mytable -a -c " ds > 'd' "Temporarily restore Cold Archive partitions without changing the storage class:
-archiveTable -t mydb.mytable -o -c " ds > 'd' "Submit a Cold Archive unarchive request without waiting:
-archiveTable -t mydb.mytable -fullTable -notWaitCheck whether Archive or Cold Archive files have been fully restored:
-archiveTable -t mydb.mytable -cr -fullTableWhat's next
To create an EMR cluster, see Create a cluster.
To use the original JindoTable archive commands (requires Jindo Namespace Service), see Use JindoTable.