After you register an E-MapReduce (EMR) cluster in DataWorks, you can customize its Kyuubi connection information. This lets you use a custom username and password to access Kyuubi and run tasks.
Background information
Apache Kyuubi is a distributed, multi-tenant gateway that provides query services such as SQL for data lake query engines such as Spark, Flink, and Trino. For more information, see Kyuubi.
Prerequisites
-
The Kyuubi service is enabled in the EMR cluster.
-
The EMR cluster is bound as a DataWorks computing resource. For more information, see Data Development (new version): Bind an EMR computing resource.
NoteWhen you bind the EMR computing resource, you must complete resource group initialization. Otherwise, you cannot find the Kyuubi configuration page.
Configure Kyuubi connection information
-
Go to the Kyuubi configuration page.
Log on to the DataWorks console. In the target region, click in the left-side navigation pane. Select a workspace from the drop-down list and click Go to Management Center.
-
In the left-side navigation pane, click Computing Resources.
-
Find the target EMR cluster and click .
-
Configure the Kyuubi connection information.
Select a connection mode based on your requirements:
-
Connection Information of Alibaba Cloud EMR Cluster: Directly log in to Kyuubi by using the Default Access Identity that you configured when you registered the EMR cluster. This mode is selected by default.
-
Custom Connection Information: If you need to use a custom username and password to log in to Kyuubi, you can select this mode. The format is
jdbc:hive2://host:port/;user=<username>;password=<password>.Note-
The first time you select Custom Connection Information, the platform automatically populates the JDBC URL based on your EMR cluster registration details; you can then modify this URL as needed.
-
If you chose to pass proxy user information when you registered the cluster, DataWorks appends the
hive.server2.proxy.userconfiguration to the JDBC URL when an EMR task runs. The rules are as follows:-
If the JDBC URL in the Custom Connection Information does not contain the placeholder
DATAWORKS_PROXY_USER, the platform appends thehive.server2.proxy.userconfiguration information to the end of the JDBC URL by default when a task is executed. -
If the JDBC URL in the Custom Connection Information contains the placeholder
DATAWORKS_PROXY_USER, the platform dynamically replaces the placeholder with the value of thehive.server2.proxy.userconfiguration when it executes a task.
-
-
-
Next steps
Follow the data development process guide to configure components for data development in DataWorks.