This topic describes how to create and manage dataflow tasks for a CPFS for Lingjun file system in the NAS console and how to view the causes of task failures.
Background information
Dataflow tasks created in the console are batch tasks. A batch task can import or export all files from one directory to another in a single run. You cannot initiate on-demand dataflows for individual files. To transfer data on a per-file basis (a streaming task), you must use OpenAPI operations. For streaming tasks, you must explicitly call the API to create a subtask for each specific file to trigger the data transfer. For more information, see Manage streaming tasks (OpenAPI).
Prerequisites
-
A dataflow has been created. For more information, see Create a dataflow within the same account or Create a dataflow across accounts.
-
The source OSS bucket must have versioning enabled. Do not suspend versioning while the dataflow is active, or the export task will fail. For more information, see Introduction to versioning.
Create a task
-
Log on to the NAS console.
-
In the left-side navigation pane, choose File System > File System List.
-
In the top navigation bar, select a region.
-
On the File System List page, click the name of the target CPFS for Lingjun file system.
-
On the details page of the file system, click Dataflow.
-
On the Dataflow page, find the target dataflow and click Task Management in the Actions column.
-
In the Task Management panel, click Create Job.
-
In the Create Job panel, select the task type and configure the parameters.
Import data
-
When a symbolic link is imported into a CPFS for Lingjun file system, it is converted into a regular file containing the link's target data, and the original link information is lost.
-
If an OSS bucket contains multiple versions of an object, only the latest version is copied.
-
File names and subdirectory names longer than 255 bytes are not supported.
-
Directory and file names cannot contain the following special characters. Otherwise, the task may produce unexpected results or fail.
-
Subdirectory or file names cannot be double periods (..).
-
Paths cannot contain backslashes (\) or consecutive backslashes (\\).
-
Subdirectory and file names cannot contain forward slashes (/).
-
-
If a file name conflicts with a subdirectory name, an object conflict occurs in the CPFS for Lingjun file system. Only one of the operations is guaranteed to succeed, and the other will fail.
Parameter
Description
Conflict Resolution Policy
The policy that specifies how to handle files with the same name in both the CPFS for Lingjun file system and the OSS bucket.
-
Skip Files with the Same Name (Default): Ignores files with the same name and does not synchronize them.
-
Keep the Latest File: Compares the modification time (mtime) of files with the same name and keeps the most recently updated version. Both OSS and CPFS for Lingjun use the modification time for comparison.
-
Overwrite Files with the Same Name: Overwrites the destination file with the version from the OSS bucket. Select Use the source file to overwrite the existing file with the same name on the destination. Make sure that you have backed up key data.
Data Type
Only the Data + Metadata type is supported. This option imports both the data blocks and metadata of files.
Specify OSS Object Prefix Subdirectory
Specifies the scope of the import task. Two modes are supported:
-
Import all files under this OSS directory: Imports all files under the specified subdirectory. Enter a relative path within the OSS Object Prefix that starts and ends with a forward slash (/).
-
Import files listed in a CSV file under this OSS directory: Imports only the files listed in a CSV file. After you select this mode, configure the CSV file list location, which is the OSS bucket directory where the CSV file is stored. The path must start and end with a forward slash (/), be 1 to 1,023 characters long, and can contain only exclamation points (!), hyphens (-), underscores (_), periods (.), and parentheses (()).
For the format requirements and an example of a CSV file list, see CSV file list format.
NoteIf the CPFS path that you configured when you created the dataflow does not exist, you can select If the CPFS directory you created does not exist, the system automatically creates a CPFS directory. to prevent import failures. Automatic directory creation is supported only by CPFS for Lingjun 2.6.0 and later.
Export data
-
The source OSS bucket must have versioning enabled. Do not suspend versioning while the dataflow is active, or the export task will fail. For more information, see Introduction to versioning.
-
When a symbolic link is synchronized to OSS, the system does not synchronize the file it points to. Instead, the symbolic link itself becomes a regular, empty object in OSS.
-
A hard link synchronizes to OSS as a regular file.
-
Files of the Socket, Device, or Pipe type cannot be exported to an OSS bucket.
-
Directory paths longer than 1,023 characters are not supported.
-
Directory and file names cannot contain the following special characters. Otherwise, the task may produce unexpected results or fail.
-
Subdirectory or file names cannot be double periods (..).
-
Paths cannot contain backslashes (\) or consecutive backslashes (\\).
-
Subdirectory and file names cannot contain forward slashes (/).
-
-
CPFS for Lingjun exports file modification timestamps to OSS custom metadata named
x-oss-meta-alihbr-sync-mtime. Do not delete or modify this metadata, or the file system timestamps will be incorrect.
Parameter
Description
Conflict Resolution Policy
The policy that specifies how to handle files with the same name in both the CPFS for Lingjun file system and the OSS bucket.
-
Skip Files with the Same Name (Default): Ignores files with the same name and does not synchronize them.
-
Keep the Latest File: Compares the modification time (mtime) of files with the same name and keeps the most recently updated version. Both OSS and CPFS for Lingjun use the modification time for comparison.
-
Overwrite Files with the Same Name: Overwrites the destination file with the version from the CPFS for Lingjun file system. Select Use the source file to overwrite the existing file with the same name on the destination. Make sure that you have backed up key data.
Export Data Type
Only the Data + Metadata type is supported. This option exports both the data blocks and metadata of files.
Specify CPFS Subdirectory
Specifies the scope of the export task. Two modes are supported:
-
Export all files under this CPFS directory: Exports all files under the specified subdirectory. Enter a relative path within the CPFS directory that starts and ends with a forward slash (/). For example,
/cpfs/. -
Export all files listed in a CSV file: Exports only the files listed in a CSV file. After you select this mode, configure the CSV file list location, which is the OSS bucket directory where the CSV file is stored. Note that in the export scenario, the CSV file is also stored in OSS. The path rules are the same as those in import mode.
For the format requirements and an example of a CSV file list, see CSV file list format.
-
-
Click OK.
Cancel a task
You can cancel a dataflow task that is running.
-
On the Dataflow tab, find the target dataflow and click Task Management in the Actions column.
-
In the Task Management panel, find the target task and click Cancel in the Actions column.
-
In the confirmation message that appears, click OK.
Copy a task
You can copy a task to run it again.
-
On the Dataflow tab, find the target dataflow and click Task Management in the Actions column.
-
In the Task Management panel, find the target task. Then, click the
icon in the Actions column and select Copy. -
In the confirmation message that appears, click OK.
CSV file list format
When you import data in CSV file list mode, each CSV file must meet the following format requirements:
-
The file must contain a Name column (case-insensitive) that specifies the relative path of each file or directory. We recommend that you also include a Size column (in bytes) to improve task splitting performance. Otherwise, the system sends a
statrequest for each file to obtain its size. -
The file name extension must be
.csvor.csv.*. Files with other extensions are ignored. -
Store the CSV files in
oss://bucket/<oss_prefix>/<csv_file_dir>/. The system reads only the files directly under this directory and does not traverse subdirectories. -
If a task uses multiple CSV files, store all of them in the same directory.
CSV file example
A path that ends with a forward slash (/) indicates a directory. Otherwise, it indicates a file. The Name column is required, and the header row in the first line must exist. Otherwise, the task returns an error.
Name,Size
/dir1/dir2/,0
/dir1/file1,1048576
/dir2/file,2097152
-
Entries that do not exist are skipped. They are not created automatically.
-
We recommend that a single CSV file contain no more than 2 million entries. If you have more entries, split them into multiple files.
-
We recommend that the total size of the CSV file list for a single dataflow task not exceed 2 GB. If it exceeds 2 GB, split the workload into multiple tasks.
-
We recommend that you sort the entries in ascending lexicographic order by path in advance for better performance.
View the cause of a task failure
When a dataflow task fails, the system either displays the cause of the failure or generates a failure report. You can view the cause or download the report in the console to troubleshoot the issue.
-
On the Dataflow tab, find the target dataflow and click Task Management in the Actions column.
-
In the Task Management panel, find the target task. Hover the pointer over the icon next to the Failed status to view the cause or download the failure report.
NoteIf no failure cause is displayed, no report is available, or you cannot resolve the issue based on the report, for assistance.
View task configuration and status
You can view the configuration and status of a batch task in the console. To view these details for a streaming task, call the DescribeDataFlowTasks API operation.
-
On the Dataflow tab, find the target dataflow and click Task Management in the Actions column.
-
In the Task Management panel, view the configuration and running status of the task.
Parameter
Description
Task ID
The unique identifier of the dataflow task.
Type
The type of the task. Valid values: Import or Export.
Conflict resolution policy
The policy for handling data with the same name that already exists in the destination file system. Valid values:
-
Skip Files with the Same Name (Default)
-
Keep the Latest File
-
Overwrite Files with the Same Name
Source address
The source, destination, and specific subdirectory paths for the data transfer.
Destination address
Source directory
Total data scanned at source
The amount of data scanned at the source, in bytes.
Data synchronized
The amount of data that has been synchronized, including skipped data, in bytes.
Data transferred
The actual amount of data transferred by the task, in bytes.
Average speed
The average transfer speed of the dataflow task. Unit: B/s.
Time remaining
The estimated time remaining until the task is complete, based on the current transfer speed.
Time period
The start and end times of the task.
Progress
The execution progress of the current task. Unit: %.
Status
The current status of the task. Valid values:
-
Pending: The dataflow task is created and is waiting in a queue.
-
Executing: The dataflow task is in progress.
-
Failed: The dataflow task failed.
-
Canceled: The dataflow task was canceled and did not complete.
-
Canceling: The dataflow task is being canceled.
-
Completed: The dataflow task is complete.
File list path
The OSS bucket directory that stores the CSV file list. This field is displayed only for tasks that use CSV file list mode. For other tasks, this field is empty.
-
View task reports
After a dataflow task is complete, the system generates a Skipped File Report, Failed File Report, or Successful File Report, depending on the outcome. You can download the report from the console to view detailed information about the files.
-
On the Dataflow tab, find the target dataflow and click Task Management in the Actions column.
-
In the Task Management panel, find the target task and click Download Task Report in the Actions column.
-
Find the report that you want to download and click the
icon.
View task performance or configure alert rules
To view task performance metrics or configure alert rules, make sure that you use a CPFS for Lingjun file system of version 2.6.0 or later and have created a dataflow task.
-
To view key performance metrics for dataflow import or export tasks, such as read/write throughput, read/write IOPS, and metadata QPS, see View CPFS performance metrics.
-
To configure alert rules for specific monitoring metrics of dataflow tasks so that you can detect and handle metric exceptions in a timely manner, see Configure a basic alert rule.