All Products
Search
Document Center

DataWorks:Flink SQL Streaming node

Last Updated:Jul 17, 2026

Flink SQL Streaming nodes in DataWorks Data Studio let you define real-time processing logic in standard SQL. They support robust state management, fault tolerance, event-time and processing-time semantics, and integrate with systems such as Kafka and HDFS. This topic describes how to develop, configure, and run a Flink SQL Streaming node to process real-time data.

Prerequisites

  • A compute resource for Realtime Compute for Apache Flink is associated in Administration. For more information, see Bind a compute engine.

  • A Flink SQL Streaming node is created. For more information, see Create a node for a scheduling workflow.

Step 1: Develop the Flink SQL Streaming node

On the Flink SQL Streaming node editing page, develop the node task.

Develop SQL code

In the SQL editor, you can define variables using the ${variable_name} format. Assign values to these variables in the Script Parameters section of the Real-Time configuration panel to pass parameters dynamically in scheduling scenarios. For example:

--Create the source table datagen_source.
CREATE TEMPORARY TABLE datagen_source(
  name VARCHAR
) WITH (
  'connector' = 'datagen'
);

--Create the result table blackhole_sink.
CREATE TEMPORARY TABLE blackhole_sink(
  name  VARCHAR
) WITH (
  'connector' = 'blackhole'
);

--Insert data from the source table into the result table.
INSERT INTO blackhole_sink
SELECT
  name
FROM datagen_source WHERE LENGTH(name) > ${name_length};
Note

In this example, the value of the parameter name_length is 5. This parameter filters the data to process only records where the name length is greater than 5 characters.

Step 2: Configure the Flink SQL Streaming node

Configure the following parameters for the Flink SQL Streaming node based on your business requirements.

Configure Flink resources

In the Flink resource information section of the Real-Time configuration panel, configure the following parameters based on the selected Resource Mode. For more information, see Configure Flink resources.

Parameter

Description

Flink cluster

The fully managed Flink compute resource associated in Administration.

Flink engine version

The engine version to use. Select a version based on your needs.

Resource Group

Select a serverless resource group that has network connectivity with Flink.

Resource Mode supports the following two modes. For more information, see Configure Flink resources.

  • Basic mode (default): Suited for beginners and simple scenarios. Uses default and simplified settings to quickly start Flink jobs.

  • Expert mode: Provides advanced options for experienced users, enabling fine-grained tuning of performance and resources for complex or high-performance scenarios.

Configure parameters based on the resource mode you selected. Understanding the Flink architecture helps you configure parameters more effectively. For more information, see Flink Architecture | Apache Flink.

Basic mode

Job Manager CPU

JobManager requires at least 0.5 CPU cores and 2 GiB of memory for stable operation. The recommended configuration is 1 CPU core and 4 GiB of memory, with a maximum of 16 CPU cores. Adjust based on cluster scale and job complexity.

Job Manager Memory

JobManager memory affects scheduling and management capacity. The recommended range is 2 GiB to 64 GiB. Adjust based on cluster scale and job requirements.

Task Manager CPU

TaskManager CPU affects task processing capability. At least 0.5 CPU cores and 2 GiB of memory are recommended, with a preferred configuration of 1 CPU core and 4 GiB of memory. The maximum is 16 CPU cores. Adjust based on your requirements.

Task Manager Memory

TaskManager memory determines the data volume and processing performance. The memory size must be at least 2 GiB and can be set up to 64 GiB.

Concurrency

The number of parallel task executions in a Flink job. Higher concurrency improves processing speed and resource utilization. Set this value based on cluster resources and job characteristics.

Number of slots per TaskManager

The number of slots per TaskManager, which determines how many tasks it can run in parallel. Adjust slots to optimize resource utilization and parallel processing.

Expert mode

Job Manager CPU

JobManager requires at least 0.25 CPU cores and 1 GiB of memory for stable operation, with a maximum of 16 CPU cores. Adjust based on cluster scale and job complexity.

Job Manager Memory

JobManager memory affects scheduling and management capacity. The recommended range is 1 GiB to 64 GiB. Adjust based on cluster scale and job requirements.

Number of slots per TaskManager

The number of slots per TaskManager, which determines how many tasks it can run in parallel. Adjust slots to optimize resource utilization and parallel processing.

Multiple SSG mode

By default, all operators share a single slot sharing group, so you cannot configure resources for each operator individually. Enable Multiple SSG mode to assign each operator its own independent slot, then configure resources on the corresponding slot.

(Optional) Configure script parameters

In the Script Parameters section of the Real-Time configuration panel in the right navigation pane, click Add parameters and edit the Parameter name and Parameter Value to dynamically use them in your code.

(Optional) Configure Flink running parameters

In the Flink running parameters section of the Real-Time configuration panel in the right navigation pane, configure the following parameters. For more information, see Configure Flink running parameters.

Parameter

Description

System Checkpoint Interval

The time interval at which Flink performs periodic checkpoints. A shorter interval reduces fault recovery time but increases system overhead. If left empty, checkpoints are disabled.

Minimum time interval between two system checkpoints

The minimum wait time between consecutive checkpoints, preventing frequent checkpoints from affecting performance. This ensures a minimum gap between two checkpoints when the maximum parallelism for checkpoints is 1.

State Data Expiration Time

The maximum time that state data can be retained without being accessed or updated. The default is 36 hours, after which state information automatically expires and is cleared to optimize storage and resource usage.

Important

This default is based on cloud best practices and differs from the open-source default (0, meaning state never expires).

Others

Additional Flink running parameters. For example: taskmanager.network.memory.max:4g.

Note

For more information about parameter configurations, see Configure Flink running parameters.

After the task is configured, click Save to save the node task.

Step 4: Start the Flink SQL Streaming node

  1. Deploy the Flink SQL Streaming node.

    The task must be deployed to Operation Center before it can run. Follow the on-screen instructions to deploy the node. For more information, see Deploy a node.

    Note

    This operation also deploys the task to the Flink VVP workspace. You can view tasks deployed through DataWorks in Flink VVP Operation Center > Job O&M.

  2. Start the Flink SQL Streaming node.

    After the task is deployed, click Go to operation and maintenance below Deploy to production environment. In Operation Center, go to Node O&M > Real-time Task O&M > Real-time computing tasks, find the task that you want to start, and click Start in the Operation column to start the real-time task and view its running status.