In a DataWorks Data Integration single-table real-time synchronization task, you can use the JSON parsing component to parse JSON data from the source into corresponding table data.
Create and configure a JSON parsing component
Step 1: Configure a Data Integration task
Create a data source. For more information, see Data source management.
Create a Data Integration task. For more information, see Configure a single-table real-time synchronization task.
NoteFor single-table real-time Data Integration tasks, you can add a data processing component between the source and destination components. For more information, see Supported data sources and synchronization solutions.
Step 2: Add a JSON parsing component
In the single-table real-time synchronization task, turn on the Data Processing switch, click +Add Node, and select the JSON Parsing component.
Enter a name and description for the node, and then configure the JSON parsing component.
ImportantTo obtain the JSON data structure, complete Data Sampling in the data source (such as Kafka) first.
Add fixed JSON parsing fields
Get formatted JSON data.
JSON source
Description
Get JSON data from data sampling
After performing data sampling, click Add Fixed Parsing Field to open the JSON Fixed Parsing Field dialog box. Select a source field and click Get JSON Structure to retrieve the JSON data structure.
Get manually entered JSON data
If you have not performed data sampling or the source data is empty, you can manually edit the fields:
Click the Edit JSON Text button to enter edit mode. In the Edit JSON Text window, manually enter the JSON content, and then click Back to Selection to select fields.
Manually Add Field. When you cannot retrieve the upstream field values and have not manually uploaded JSON content by clicking the Edit JSON Text button, you can manually define fixed field parsing rules by editing the value content to obtain JSON content. The parameters for manually added fields are as follows:
Parameter
Description
Field
The reference name of the parsed new field in downstream nodes.
Value acquisition
Specifies the JSON parsing path. The parsing syntax is as follows:
$: represents the root node..: represents a child node.[]:[number]represents an array index, starting from 0.[*]: expands the array into multi-row output. Each element is combined with other fields of the record to form an independent row for output to downstream nodes.
NoteOnly letters, digits, hyphens (-), and underscores (_) are allowed in JSON field names within the JSON parsing path.
Default
The default value when the JSON parsing path does not exist due to changes in upstream table fields.
NULL: The field is assigned the value NULL.
Do Not Fill: The field is not filled with any value. The difference from NULL is that when writing to the corresponding field in the destination table, if the destination field has a configured default value, that default value is used instead of NULL.
Dirty Data: The record is counted in the dirty data statistics of the synchronization task, and whether the task exits abnormally is determined by the dirty data tolerance configuration.
Manually enter a constant: Use a manually entered constant as the field value.
Add dynamic JSON parsing fields
In the formatted JSON content, select the target JSON object field for dynamic parsing. The system automatically adds a parsing configuration for each field under the JSON object in the fixed output fields.
After you set dynamic parsing for a JSON object, during actual synchronization task execution, each field under the specified path in the JSON object is added to the record as a STRING type with its original JSON field name and value, and output to downstream nodes. This way, when the structure of the specified JSON object changes or new fields are added during synchronization, they are automatically identified and output to downstream nodes.
Get formatted JSON data.
JSON source
Description
Get JSON data from data sampling
After performing data sampling, click Add Dynamic Parsing Field to open the JSON Dynamic Output Fields dialog box. Select a source field and click Get JSON Structure to retrieve the JSON data structure.
Get manually entered JSON data
If you have not performed data sampling or the source data is empty, you can manually edit the fields:
Click the Edit JSON Text button to enter edit mode. In the Edit JSON Text window, manually enter the JSON content, and then click Back to Selection to select fields.
Dynamic parsing of JSON objects
Configuration method:
In the left-side JSON Data Structure panel, click the select button next to the target JSON object field (such as dynamic). A Specify JSON Object configuration is automatically added to the right-side Dynamic Output Fields section, with the Value automatically populated with the corresponding path (such as
$.dynamic) and the Default Value set to Ignore.Assume that the dynamic field adds c3. The parsing results before and after the change are as follows:
_value_(STRING)
c1(STRING)
c2(STRING)
c3(STRING)
{ "dynamic": { "c1": 2, "c2": ["a1","b1"] } }2["a1","b1"]Not filled
{ "dynamic": { "c1": 2, "c2": ["a1","b1"], "c3": {"name": "jack"} } }2["a1","b1"]{"name": "jack"}
Manually add fields.
Manually adding fields refers to manually defining dynamic field parsing rules by editing the value content when you cannot retrieve the upstream subfield values and have not manually uploaded JSON content by clicking the Edit JSON Text button:
Parameter
Description
Specify JSON Object
Specifies the JSON object parsing path. The parsing syntax is as follows:
$: represents the root node..: represents a child node.[]: [number] represents an array index, starting from 0.
Note: Only letters, digits, hyphens (-), and underscores (_) are allowed in JSON field names within the JSON parsing path.
Default
Specifies the default behavior when the specified JSON parsing path fails to parse or the corresponding field does not exist.
Ignore: Do not perform dynamic parsing.Dirty Data: The record is counted in the dirty data statistics of the synchronization task, and whether the task exits abnormally is determined by the dirty data tolerance configuration.
Policy when a duplicate field name is found.
When JSON dynamic fields are expanded by key-value, only the first level is expanded. If a field with the same name already exists during expansion, the following policies are applied:
Overwrite: The value from the later expansion replaces the existing field value.
Discard: The existing field value is retained, and the value from the later expansion is discarded.
Error: The task reports an error and stops running.
Next step
After you configure the Data Source and JSON Parsing components, click Output Preview to check whether the output data of the current node meets your requirements.