All Products
Search
Document Center

E-MapReduce:Data source management

Last Updated:Aug 20, 2026

You can use the data source management features of E-MapReduce (EMR) Workflow to configure data sources to meet your different requirements for data storage and access. This topic describes how to create, modify, and delete a data source.

Limits

The cluster in which the data source resides and the cluster used to run a workflow must be deployed in the same virtual private cloud (VPC).

Create a data source

  1. Go to the Datasource page.

    1. Log on to the EMR console using your Alibaba Cloud account or a RAM user.

    2. In the left-side navigation pane, choose EMR Workbench > Workflow.

    3. On the Workflow page, find the target workspace and click Console in the Actions column.

    4. Click the Data Source Center tab.

  2. On the Data Source Center page, click Create Data Source.

  3. In the CreateDataSource dialog box, configure the parameters. The following table describes the parameters.

    Hive/Impala data source

    Parameter

    Required

    Description

    Datasource

    Yes

    The type of the data source.

    Data Source Name

    Yes

    The name of the data source.

    Description

    No

    The description of the data source.

    Host IP Address

    Yes

    The IP address or hostname for the Hive/Impala connection.

    Port

    Yes

    The port for the Hive/Impala data source is 10000.

    Username

    Yes

    The username for the Hive/Impala connection.

    Password

    No

    The password for the Hive/Impala connection.

    Database Name

    Yes

    The database name for the Hive/Impala connection.

    JDBC connection parameters

    No

    The parameters for the data source connection. The format is {"key1":"value1","key2":"value2"...}.

    Test Connectivity

    No

    You can use a scheduling resource group to test the connection when you add the data source.

    Note
    • If a workflow uses this data source, ensure connectivity between the data source and the scheduling resource group.

    • You can only test connectivity between the data source and the default resource group or cluster resource groups.

    Presto data source

    Parameter

    Required

    Description

    Datasource

    Yes

    The type of the data source.

    Data Source Name

    Yes

    The name of the data source.

    Description

    No

    The description of the data source.

    Host IP Address

    Yes

    The IP address or hostname for the data source connection.

    Port

    Yes

    The port for the Presto data source is 8080.

    Username

    Yes

    The username for the Presto connection.

    Password

    No

    The password for the Presto connection.

    Catalog

    No

    The name of the catalog for the Presto connection.

    Database Name

    Yes

    The name of the database for the Presto connection.

    JDBC connection parameters

    No

    The parameters for the data source connection. The format is {"key1":"value1","key2":"value2"...}.

    Test Connectivity

    No

    You can use a scheduling resource group to test the connection when you add the data source.

    Note
    • If a workflow uses this data source, ensure connectivity between the data source and the scheduling resource group.

    • You can only test connectivity between the data source and the default resource group or cluster resource groups.

    Doris data source

    Parameter

    Required

    Description

    Datasource

    Yes

    The type of the data source.

    Data Source Name

    Yes

    The name of the data source.

    Description

    No

    The description of the data source.

    Host IP Address

    Yes

    The IP address or hostname for the Doris connection.

    Port

    Yes

    The port for the Doris data source is 9030.

    Username

    Yes

    The username for the Doris connection.

    Password

    No

    The password for the Doris connection.

    FE Endpoint

    No

    The IP address and port of the frontend (FE) node, in ip:port format. Use commas to separate multiple nodes.

    Database Name

    Yes

    The name of the database for the Doris connection.

    JDBC connection parameters

    No

    The parameters for the Doris connection. The format is {"key1":"value1","key2":"value2"...}.

    Test Connectivity

    No

    You can use a scheduling resource group to test the connection when you add the data source.

    Note
    • If a workflow uses this data source, ensure connectivity between the data source and the scheduling resource group.

    • You can only test connectivity between the data source and the default resource group or cluster resource groups.

    SSH data source

    Parameter

    Required

    Description

    Datasource

    Yes

    The type of the data source.

    Data Source Name

    Yes

    The name of the data source.

    Description

    No

    The description of the data source.

    Host IP Address

    Yes

    The IP address or hostname for the SSH connection.

    Port

    Yes

    The port for the SSH data source is 22.

    Username

    Yes

    The username for the SSH connection.

    Password

    No

    The password for the SSH connection.

    Private key

    No

    The private key for the SSH connection.

    Test Connectivity

    No

    You can use a scheduling resource group to test the connection when you add the data source.

    Note
    • If a workflow uses this data source, ensure connectivity between the data source and the scheduling resource group.

    • You can only test connectivity between the data source and the default resource group or cluster resource groups.

    StarRocks data source

    Parameter

    Required

    Description

    Datasource

    Yes

    The type of the data source.

    Data Source Name

    Yes

    The name of the data source.

    Description

    No

    The description of the data source.

    Host IP Address

    Yes

    The IP address or hostname for the StarRocks connection.

    Port

    Yes

    The port for the StarRocks data source is 9030.

    Username

    Yes

    The username for the StarRocks connection.

    Password

    No

    The password for the StarRocks connection.

    FE Endpoint

    No

    The IP address and port of the FE node. The format is ip:port. Use commas to separate multiple nodes, for example: ip1:port1,ip2:port2.

    Database Name

    Yes

    The name of the database for the StarRocks connection.

    JDBC connection parameters

    No

    The parameters for the StarRocks connection. The format is {"key1":"value1","key2":"value2"...}.

    Test Connectivity

    No

    You can use a scheduling resource group to test the connection when you add the data source.

    Note
    • If a workflow uses this data source, ensure connectivity between the data source and the scheduling resource group.

    • You can only test connectivity between the data source and the default resource group or cluster resource groups.

  4. Click OK.