All Products
Search
Document Center

DataWorks:ADB Spark node

Last Updated:Aug 20, 2026

Use an ADB Spark node in DataWorks to develop, periodically schedule, and integrate AnalyticDB Spark tasks with other jobs. This topic describes the main workflow for developing tasks by using an ADB Spark node.

Background information

ADB Spark is a compute engine within the AnalyticDB service designed to run large-scale Apache Spark data processing tasks. It supports real-time data analysis, complex queries, and machine learning applications. With support for Java, Scala, and Python, it simplifies development and automatically scales to optimize performance and costs. You can configure tasks by uploading relevant Jar or .py files. ADB Spark is ideal for various industries that need to efficiently process massive datasets and gain real-time insights. This helps enterprises extract valuable information from data to drive business growth.

Prerequisites

AnalyticDB for MySQL prerequisites:

  • You have created an AnalyticDB for MySQL Basic Edition cluster in the same region as your DataWorks workspace. For more information, see Create a cluster.

  • You have configured a job resource group in your AnalyticDB for MySQL cluster. For more information, see Create a job resource group.

    Note

    When you use DataWorks to develop Spark applications, you must create a job resource group.

  • If you use OSS for storage, ensure the OSS bucket is in the same region as your AnalyticDB for MySQL cluster.

DataWorks prerequisites:

  • You have a workspace with the Use Data Studio (New Version) option selected and a resource group attached. For more information, see Create a workspace.

  • The resource group must be in the same VPC as the AnalyticDB for MySQL cluster, and you must add the IP addresses of the resource group to the whitelist of the AnalyticDB for MySQL cluster. For more information, see Configure a whitelist.

  • You have added your AnalyticDB for MySQL cluster instance to DataWorks as a compute engine. The compute engine type must be AnalyticDB for Spark. DataWorks tests the compute engine connectivity using the resource group. For more information, see Bind a compute engine.

  • You have created a workflow folder. For more information, see Workflow folder.

  • You have created an ADB Spark node. For more information, see Create a node for a workflow.

Step 1: Develop the ADB Spark node

In the ADB Spark node editor, you can configure the node content based on the selected Language type by using the sample spark-examples_2.12-3.2.0.jar package or the spark_oss.py file. For more information about how to develop the node content, see Develop a Spark application by using the spark-submit command-line tool.

Java and Scala

Prepare the JAR file

You must upload the sample JAR package to OSS. The node requires this package to execute the task.

  1. Prepare the sample JAR package.

    Download the spark-examples_2.12-3.2.0.jar sample JAR package for use in the ADB Spark node.

  2. Upload the sample code to OSS.

    1. Log on to the OSS console. In the left-side navigation pane, click Buckets.

    2. On the Buckets page, click Create Bucket. In the Create Bucket panel, create a bucket in the same region as your AnalyticDB for MySQL cluster.

      Note

      This topic uses a bucket named dw-1127 as an example.

    3. Create an external storage directory.

      After the creation is complete, click Go to Bucket. On the Files page, click Create Directory, and set Catalog Name to db_home.

    4. Upload the sample code file spark-examples_2.12-3.2.0.jar to the db_home directory. For more information, see Upload objects.

Configure the ADB Spark node

Configure the ADB Spark node content based on the following parameter descriptions.

Language

Parameter

Description

Java/Scala

Main JAR Resource

The OSS path to the JAR package. Example: oss://dw-1127/db_home/spark-examples_2.12-3.2.0.jar.

Main Class

The main class in your compiled JAR package. The name of the main class in the sample code is org.apache.spark.examples.SparkPi.

Parameter

The arguments to pass to your code. You can configure this field with dynamic parameters in the ${var} format.

Note

In this example, the dynamic parameter ${var} can be set to 1000.

Configuration Item

The runtime parameters for the Spark program. For more information, see Spark application configuration parameters. Example:

spark.driver.resourceSpec:medium

Python

Prepare the Python file and data

Follow these steps to upload the test data file and sample code to OSS. This enables the sample code in the node configuration to read the test data file.

  1. Prepare the test data.

    Create a file named data.txt and add the following content.

    Hello,Dataworks
    Hello,OSS
  2. Write the sample code.

    Create a file named spark_oss.py and add the following content to the spark_oss.py file.

    import sys
    
    from pyspark.sql import SparkSession
    
    # Initialize Spark.
    spark = SparkSession.builder.appName('OSS Example').getOrCreate()
    # Read the specified file. The file path is specified by the value of the argument passed to the job.
    textFile = spark.sparkContext.textFile(sys.argv[1])
    # Calculate and print the number of lines in the file.
    print("File total lines: " + str(textFile.count()))
    # Print the first line of the file.
    print("First line is: " + textFile.first())
    
  3. Upload the test data and sample code to OSS.

    1. Log on to the OSS console. In the left-side navigation pane, click Buckets.

    2. On the Buckets page, click Create Bucket. In the Create Bucket panel, create a bucket in the same region as your AnalyticDB for MySQL cluster.

      Note

      This topic uses a bucket named dw-1127 as an example.

    3. Create an external storage directory.

      After the creation is complete, click Enter Bucket. On the Files page, click Create Directory, and set the Catalog Name to db_home.

    4. Upload the test data file data.txt and the sample code file spark_oss.py to the db_home directory. For more information, see Upload objects.

Configure the ADB Spark node

Configure the ADB Spark node content based on the following parameter descriptions.

Language

Parameter

Description

Python

Main Package

The path to the sample code file. Example: oss://dw-1127/db_home/spark_oss.py.

Parameter

The arguments to pass to your code. For this example, use the storage path of the test data file. Example: oss://dw-1127/db_home/data.txt.

Configuration Item

The runtime parameters for the Spark program. For more information, see Spark application configuration parameters. Example:

spark.driver.resourceSpec:medium

Step 2: Debug the ADB Spark node

  1. Configure the debug properties for the ADB Spark node.

    In the Run Configuration pane on the right side of the node editor, configure the Compute Resource, ADB Computing Resource Group, Resource Group, and Compute Units parameters.

    Parameter type

    Parameter

    Description

    Compute Resource

    Compute Resource

    Select the attached AnalyticDB for Spark compute engine.

    ADB Computing Resource Group

    Select the job resource group that you created in the AnalyticDB for MySQL cluster. For more information, see Resource group overview.

    Resource Group

    Resource Group

    Select the resource group that passed the connectivity test when you attached the compute engine.

    Compute Units

    This node uses the default number of CUs. You do not need to modify this value.

  2. Debug and run the ADB Spark node.

    To run the node task, click Save and then Run.

Step 3: Schedule the ADB Spark node

  1. Configure the scheduling properties for the ADB Spark node.

    To run the node task periodically, open the Scheduling Settings pane on the right side of the node editor. In the Scheduling Policy section, configure the following parameters as needed. For more information about other parameters, see Configure scheduling for a node.

    Parameter

    Description

    Compute Resource

    Select the attached AnalyticDB for Spark compute engine.

    ADB Computing Resource Group

    Select the job resource group that you created in the AnalyticDB for MySQL cluster. For more information, see Resource group overview.

    Resource Group

    Select the resource group that passed the connectivity test when you attached the compute engine.

    Compute Units

    This node uses the default number of CUs. You do not need to modify this value.

  2. Publish the ADB Spark node.

    After you configure the node task, you must publish the node. For more information, see Publish a node or workflow.

Next steps

After publishing the task, you can view its running status in Operation Center. For more information, see Getting started with Operation Center.