Introduction to one-click data backfill
Running backfill tasks manually is time-consuming and error-prone. The one-click data backfill feature automates this process by expanding all tasks in the correct order and completing the data backfill. If an error occurs, you can go to the DataWorks Operation Center to view the details.
Instructions
After the recommendation algorithm code is deployed to DataWorks, click to view the task backfill process, and then click to create a backfill task.
In the dialog box, modify the task end date and click OK to create the backfill task. The default maximum concurrency is 2. When you backfill feature tasks for multiple days, setting this value to 2 or 4 helps the task complete faster. The maximum run time is measured in hours. To avoid affecting production tasks that run early the next morning, set a maximum run time as needed.
Click the backfill task list, and then click to start the tasks sequentially to run the backfill tasks.
For a failed task, click to view the task node. This redirects you to DataWorks to view the task and troubleshoot the problem.
If you need to regenerate the code and rerun a backfill task, set the feature and model versions when you configure the custom algorithm. This prevents data overwrites and conflicts.
This feature automates backfill tasks from a start date to an end date. If recent data is missing from a completed backfill task, you can use the Auto Triggered Nodes feature in the DataWorks Operation Center to perform an incremental backfill.