The MoveTo command migrates table and partition data to a new storage path.
Prerequisites
- Java Development Kit (JDK) 8 is installed.
- A cluster is created. For more information, see Create a cluster.
Background information
The MoveTo command copies underlying data and automatically updates metadata to fully migrate table and partition data to a new path. You can use filter conditions to copy multiple partitions at a time. Built-in safeguards protect data integrity during migration.
Limits
The MoveTo command is supported on clusters that run EMR V3.36.0 or later, or EMR V5.2.0 or later.
Use the MoveTo command
Important Only one MoveTo process can run on a cluster at a time. If a MoveTo process is already running, a new MoveTo process fails to obtain the configuration lock and exits. The system notifies you that a process is running. In this case, you can stop the running process and start a new one, or wait for the current process to finish.
- Log on to the cluster using Secure Shell (SSH). For more information, see Log on to a cluster.
- Run the following command to view the help information.
jindo table -help moveToThe help information is similar to the following:<dbName.tableName> The table to move. <destination path> The destination base directory which is always at the same level of a 'table location', where the moved partitions or un-partitioned data would located in. <condition>/-fullTable A filter condition to determine which partitions should be moved, supporting common operators (like '>') and built-in UDFs (like to_date) (UDFs not supported yet...), while -fullTable means that all partitions (or a whole un-partitioned table) should be moved. One but only one option must be specified among -c "<condition>" and -fullTable. <before days> Optional, saying that table/partitions should be moved only when they are created (not updated or modified) more than some days before from now. <parallelism> The maximum concurrency when copying partitions, 1 by default. <OSS storage policy>: Storage policy for OSS destination, which can be Standard (by default), IA, Archive, or ColdArchive. Not applicable for destinations other than OSS. NOTE: if you are willing to use ColdArchive storage policy, please make sure that Cold Archive has been enabled for your OSS bucket. -o/-overWrite Overwriting the final paths where the data would be moved. For partitioned tables this overwrites partitions' locations which are subdirectories of <destination path>; for un-partitioned table this overwrites the <destination path> itself. -r/-removeSource Let the source data be removed when the corresponding table/partition is successfully moved to the new destination. Otherwise (by default), the source data would be left as it was. -skipTrash Applicable only when [-r/-removeSource] is enabled. If present, source data would be immediately deleted from the file system, bypassing the trash. -e/-explain If present, the command would not really move data, but only prints the table/partitions that would be moved for given conditions. <log directory> A directory to locate log files, '/tmp/<current user>/' by default.The syntax of the MoveTo command is as follows:jindo table -moveTo \ -t <dbName.tableName> \ -d <destination path> \ [-c "<condition>" | -fullTable] \ [-b/-before <before days>] \ [-p/-parallel <parallelism>] \ [-s/-storagePolicy <OSS storage policy>] \ [-o/-overWrite] \ [-r/-removeSource] \ [-skipTrash] \ [-e/-explain] \ [-l/-logDir <log directory>]Parameter Description Required -t <dbName.tableName> The name of the table to move. Use the format database_name.table_name.The database name and table name are separated by a period (.). The table can be partitioned or non-partitioned.
Yes -d <destination path> The destination path. Whether you move partitions or an entire non-partitioned table, this path corresponds to the table-level location. If you move a partition, the full path of the partition is the destination path plus the partition name, such as <destination path>/p1=v1/p2=v2/.Yes -c "<condition>" | -fullTable You must specify either -c "<condition>"or-fullTable.- If you specify
-fullTable, the entire table is moved. The table can be partitioned or non-partitioned. - If you specify
-c "<condition>", a filter condition is provided to select the partitions to move. Common operators, such as the greater-than sign (>), are supported.For example, for a partition key column `ds` of the String data type, if you want to select partitions where the partition name is greater than 'd', use
-c " ds > 'd' ".
No -b/before <before days> Moves only tables or partitions that were created more than a specified number of days ago. No -p/-parallel <parallelism> The degree of parallelism for the migration operation. No -s/-storagePolicy <OSS storage policy> The storage policy for data that is copied to OSS. The following policies are available: - Standard: (Archive Storage).
- IA: Infrequent Access (IA) storage.
- Storage class: Standard
- ColdArchive: Cold Archive storage.Note Make sure that this feature is enabled for your OSS bucket before you use this policy.
No -o/-overWrite Specifies whether to overwrite the destination path. For a partitioned table, this option clears only the paths of the partitions to be moved, not the entire table path. No -r/-removeSource Specifies whether to remove the source data after the migration is complete and the metadata is updated. For a partitioned table, this option removes only the source paths of the partitions that were successfully moved. No -skipTrash Specifies whether to bypass the Trash when the source path is removed. Note This parameter is applicable only when -r/-removeSource is specified.No -e/-explain If this option is specified, the command runs in explain mode. It lists the partitions that would be moved but does not actually move any data. No -l/-logDir <log directory> Specifies the directory for log files. No - If you specify
Configure the lock directory
The MoveTo tool uses a process lock, which requires a Hadoop Distributed File System (HDFS) path to store the LOCK file. By default, the path is hdfs:///tmp/jindotable-lock/.
Important The path for the LOCK file must be an HDFS path. If you do not have the required permissions for the default path, follow these steps to configure a custom path.
- Go to the HDFS service page.
- Log on to the Alibaba Cloud E-MapReduce console.
- In the top menu bar, select the region and resource group where your cluster resides.
- Click the Clusters tab.
- On the Clusters page, click Details in the row of the desired cluster.
- In the navigation pane on the left, choose .
- Log on to the Alibaba Cloud E-MapReduce console.
- Modify the configuration.
- On the HDFS service page, click the Configure tab. Then, click the hdfs-site or core-site tab.
- In the upper-right corner, click Custom Configuration.
- In the Add Configuration Item dialog box, add the jindotable.moveto.tablelock.base.dir configuration item and set its value to an existing HDFS path.
Important Before you configure a custom lock directory, make sure that no MoveTo process is running on any node in the cluster. Otherwise, the MoveTo operation may fail or cause data corruption.
- Save the configuration.
- In the upper-right corner, click Save.
- In the Confirm dialog box, enter a reason for the change and click OK.