MaxCompute supports connecting to Hadoop clusters by creating external data sources, which lets you build a lakehouse architecture. This topic describes how to create, view, and delete a Hadoop external data source.
Background information
A Hadoop external data source maps to an external MaxCompute project, so you can run single-source or federated queries on the Hadoop data in MaxCompute. MaxCompute supports creating, viewing, and deleting Hadoop external data sources.
Usage notes
-
Only the China (Hangzhou), China (Shanghai), China (Beijing), China (Shenzhen), China (Zhangjiakou), and Singapore regions support creating external data sources.
-
An external data source can be bound to only one external MaxCompute project. Multiple external projects cannot bind the same external data source.
-
External data sources support creation, viewing, and deletion only. You cannot update an external data source.
Create an external data source
-
Log on to the MaxCompute console and select the region you want.
-
On the External Data Source Management tab, click Create External Data Source Link.
-
In the Create External Data Source dialog box, configure the parameters described in the following table, and then click OK.
Parameter
Description
Select MaxCompute Project
Select the destination MaxCompute project. You can view MaxCompute project names on the Projects tab.
External Data Source Name
A custom name for the external data source. The naming rules are as follows:
The name can contain only lowercase letters, digits, and underscores.
The name must be fewer than 128 characters.
Network Connection Object
The connection from MaxCompute to the E-MapReduce or Hadoop VPC. Configure the connection based on VPC connection scheme.
NameNode Address
The service addresses and port numbers of the active and standby NameNode processes of the destination Hadoop cluster. The port number is usually
8020. Contact your Hadoop cluster administrator for the exact values.HMS Service Address
The addresses and port numbers of the Hive Metastore Service (HMS) on the active and standby NameNodes of the destination Hadoop cluster. The port number is usually
9083. Contact your Hadoop cluster administrator for the exact values.Cluster Name
The name that identifies the NameNode in a high availability (HA) Hadoop cluster. For a self-managed Hadoop cluster, you can get the name from the
dfs.nameservicesparameter in thehdfs-site.xmlfile.Authentication Type
MaxCompute uses account mapping to obtain metadata and data from the Hadoop cluster. The mapped Hadoop account is usually protected by an authentication and authorization mechanism such as Kerberos, so the authentication and authorization file of the account is required. Select the type that matches your cluster, and consult your Hadoop O&M engineer for details.
-
No Authentication Method: Select this option if Kerberos authentication is disabled for the Hadoop cluster.
-
Kerberos Account Authentication: Select this option if Kerberos authentication is enabled for the Hadoop cluster.
-
Configuration File: Upload the
krb5.conffile of the Hadoop cluster.NoteIf the Hadoop cluster runs on a Linux operating system, the
krb5.conffile is usually in the/etcdirectory on the master node that hosts the Hadoop HDFS NameNode. -
hmsPrincipals: The identity of the HMS service. Run the
list_principalscommand on the Kerberos terminal of the Hadoop cluster to get the HMS principals. An example value is as follows.hive/emr-header-1.cluster-20****@EMR.20****.COM,hive/emr-header-2.cluster-20****@EMR.20****.COMNoteThe service information of different nodes is a comma-delimited string, and each principal maps to one HMS service address.
-
Add Configuration Engine Permission Mapping.
-
Cloud Account: The Alibaba Cloud account that MaxCompute uses to access the Hadoop cluster.
-
Kerberos Account: The Kerberos-authorized Hadoop user account that has Hive access permissions.
-
Upload: Upload the keytab configuration file of the Kerberos account. You can generate the file as described in Create a keytab configuration file.
-
-
View or delete an external data source
-
Log on to the MaxCompute console and select the region you want.
-
On the External Data Source Management tab of the MaxCompute console, click Details or Delete in the Actions column of the external data source you want to manage.
If the external data source is already bound to an External Project, you cannot delete it. Delete the External Project or unbind it from the external data source first, and then delete the external data source.