You can configure metadata for a Spark cluster in EMR on ACK by using either Data Lake Formation (DLF) or a self-managed Hive metastore. This topic describes how to do this.
Background information
Data Lake Formation (DLF) provides high availability and is easy to maintain, making it a suitable choice for the following use cases:
-
When all your EMR clusters are in a production environment, you do not need to maintain a separate metadatabase.
-
When you use multiple big data compute engines such as MaxCompute, Hologres, and Platform for AI, you can manage metadata centrally.
-
When you have multiple EMR clusters, you can manage their metadata uniformly.
Prerequisites
-
You have created a Spark cluster in the EMR on ACK console. For more information, see Step 1: Create a cluster.
-
If you want to use Data Lake Formation (DLF), ensure DLF is activated. For more information, see Quick start.
-
If you want to use a self-managed Hive metastore, you must have a Hive metastore service that is network-accessible from your ACK cluster.
Option 1: Use Data Lake Formation (DLF) (recommended)
-
Go to the cluster details page.
-
Log on to the EMR on ACK console.
-
On the EMR on ACK page, click the name of the target cluster.
-
-
On the Cluster Details page, click Enable next to Data Lake Formation (DLF).
-
In the Enable DLF dialog box, click OK.
After you complete this configuration, jobs that you submit to the Spark cluster automatically connect to DLF.
Option 2: Use a self-managed Hive metastore
-
Go to the cluster's configuration page.
-
Log on to the EMR on ACK console.
-
On the EMR on ACK page, find the target cluster and click Configure in the Actions column.
-
-
On the Configure tab, click the spark-defaults.conf tab.
-
Add a custom configuration.
-
At the top of the page, click Add Configuration Item.
-
Set the Key to spark.hadoop.hive.metastore.uris and the Value to thrift://<IP address of your self-managed Hive metastore>:9083.
This parameter specifies the URI for connecting to the Hive metastore by using the Thrift protocol. Modify this value to match your environment.
-
Click OK.
-
In the dialog box that appears, enter a reason and click Save.
-
-
Deploy the client configuration.
-
Click Deploy Client Configuration.
-
In the dialog box that appears, enter a reason and click OK.
-
In the Confirm dialog box, click OK.
After you complete this configuration, jobs submitted to the Spark cluster automatically connect to your self-managed Hive metastore.
-