All Products
Search
Document Center

E-MapReduce:Archive and unarchive data in SDK mode

Last Updated:Sep 16, 2026

Use JindoTable's archiveTable and unarchiveTable commands to move Hive tables and partitions between OSS storage classes — without requiring the Jindo Namespace Service component of SmartData.

Prerequisites

Before you begin, ensure that you have:

  • An E-MapReduce (EMR) cluster running EMR V3.36.0 or later, or EMR V5.2.0 or later

  • A partitioned or non-partitioned table stored in OSS (only table data can be archived)

  • SSH access to the cluster

Background

The original JindoTable archive and unarchive commands depend on the Jindo Namespace Service component of SmartData. The archiveTable and unarchiveTable commands remove that dependency, which means:

  • They work on clusters where SmartData is not deployed, including self-managed clusters.

  • They support filter parameters for multi-threaded archiving across large partition sets. If local multithreading is not enough, use MapReduce to distribute the work across the entire cluster.

For documentation on the original commands, see Use JindoTable.

Storage class transitions

The following table shows which storage classes each command can target and which files are skipped during the operation.

archiveTable

FlagTarget storage classFiles skipped
-iInfrequent Access (IA)Archive, Cold Archive
-aArchiveCold Archive
-caCold ArchiveNone

unarchiveTable

FlagResulting storage classFiles skipped
(none)StandardNone
-iInfrequent Access (IA)Standard
-aArchiveStandard, IA
-oUnchanged (temporary restore only)Standard, IA, previously restored files
-crNo change (check restore status only)N/A
Note The storage class of Cold Archive data can only change to Standard, IA, or Archive after the data is completely unarchived.

Archive tables or partitions

Use archiveTable to move OSS table data to a lower-cost storage class.

  1. Log in to your cluster over SSH. See Log on to a cluster.

  2. Run the following command to see full help:

    jindo table -help archiveTable
  3. Archive tables or partitions using the following syntax:

    -archiveTable -t <dbName.tableName> \
    -i/-a/-ca \
    [-c "<condition>" | -fullTable] \
    [-b/-before <before days>] \
    [-p/-parallel <parallelism>] \
    [-mr/-mapReduce] \
    [-e/-explain] \
    [-w/-workingDir <working directory>] \
    [-l/-logDir <log directory>]

Parameters

ParameterDescriptionRequired
-t <dbName.tableName>The table to archive, in database_name.table_name format. Supports partitioned and non-partitioned tables.Yes
-i / -a / -caTarget storage class: -i for Infrequent Access (IA), -a for Archive, -ca for Cold Archive. See Storage class transitions for skip rules.Yes
-c "<condition>" / -fullTableScope of archiving: use -fullTable for the entire table, or -c "<condition>" to archive only matching partitions. Supports standard comparison operators.Yes
-b / -before <before days>Only archive tables or partitions created at least N days ago.No
-p / -parallel <parallelism>Number of parallel archiving threads.No
-mr / -mapReduceUse Hadoop MapReduce instead of local multithreading.No
-e / -explainDry run: list partitions that would be archived without making any changes.No
-w / -workingDir <directory>Working directory for the MapReduce job. Must have read and write permissions. Temporary files are created during the job and deleted automatically on completion.No
-l / -logDir <directory>Directory for log files.No

Examples

Preview which partitions will be archived (dry run):

-archiveTable -t mydb.mytable -a -c " ds > 'd' " -e

Archive all partitions in a table to Cold Archive:

-archiveTable -t mydb.mytable -ca -fullTable

Archive partitions where the `ds` column is greater than `'d'`, created more than 30 days ago:

-archiveTable -t mydb.mytable -a -c " ds > 'd' " -b 30

Unarchive tables or partitions

Use unarchiveTable to restore OSS table data from archived storage. Its syntax mirrors archiveTable, with two differences: the storage class flag (-i/-a/-o/-cr) is optional, and the -notWait flag is available for Cold Archive restores.

  1. Log in to your cluster over SSH. See Log on to a cluster.

  2. Run the following command to see full help:

    jindo table -help unarchiveTable
  3. Unarchive tables or partitions using the following syntax:

    -archiveTable -t <dbName.tableName> \
    [-i/-a/-o/-cr] \
    [-notWait] \
    [-c "<condition>" | -fullTable] \
    [-b/-before <before days>] \
    [-p/-parallel <parallelism>] \
    [-mr/-mapReduce] \
    [-e/-explain] \
    [-w/-workingDir <working directory>] \
    [-l/-logDir <log directory>]

Parameters unique to unarchiveTable

All parameters from archiveTable apply. The following parameters differ or are new:

ParameterDescriptionRequired
-i / -a / -o / -crTarget storage class after unarchiving. If omitted, data is restored to Standard. See Storage class transitions for details.No
-notWaitValid only when unarchiving data. Submit the unarchive request for Cold Archive data and return immediately without waiting for completion.No

Examples

Preview which partitions will be restored (dry run):

-archiveTable -t mydb.mytable -a -c " ds > 'd' " -e

Restore all partitions to Standard storage:

-archiveTable -t mydb.mytable -fullTable

Restore partitions to Archive storage, skipping Standard and IA files:

-archiveTable -t mydb.mytable -a -c " ds > 'd' "

Temporarily restore Cold Archive partitions without changing the storage class:

-archiveTable -t mydb.mytable -o -c " ds > 'd' "

Submit a Cold Archive unarchive request without waiting:

-archiveTable -t mydb.mytable -fullTable -notWait

Check whether Archive or Cold Archive files have been fully restored:

-archiveTable -t mydb.mytable -cr -fullTable

What's next

  • To create an EMR cluster, see Create a cluster.

  • To use the original JindoTable archive commands (requires Jindo Namespace Service), see Use JindoTable.