The E-MapReduce Hook collects data lineage and access frequency from SparkSQL jobs. You can use Data Lake Formation (DLF) to track table and partition access counts, or DataWorks to manage data lineage.
Prerequisites
A DataLake or custom cluster with the Spark service is required. For more information, see Create a cluster.
Limitations
-
EMR-HOOK does not support collecting SQL information from jobs in a gateway environment custom-deployed with EMR-CLI.
-
In EMR versions earlier than 5.16.0 and 3.50.0, the hive.exec.post.hooks (for Hive) and spark.sql.queryExecutionListeners (for Spark) parameter settings are not synchronized to gateway nodes. In EMR versions 5.16.0 and later and 3.50.0 and later, these settings are synchronized to gateway nodes. These versions also introduce the hive_aux_jars_path_gateway_only parameter, which lets you independently use custom extension JAR files on gateway nodes.
Usage notes
-
In E-MapReduce versions earlier than EMR-5.14.0 and EMR-3.48.0, the E-MapReduce Hook is enabled by default.
You can manually enable the E-MapReduce Hook on custom clusters that run EMR-3.44, as it is disabled by default. For more information, see FAQ.
-
In E-MapReduce versions EMR-5.14.0 and later, and EMR-3.48.0 and later, the E-MapReduce Hook is disabled by default, and you must enable it manually.
Procedure
-
Go to your cluster's Services page.
-
Log on to the E-MapReduce console.
-
In the top navigation bar, select a region and a resource group based on your requirements.
-
On the EMR on ECS page, find the target cluster and click Services in the Services column.
-
-
Configure the E-MapReduce Hook.
-
On the Services page, find the Spark2 or Spark3 service and click Configure.
-
On the Configure page, on the corresponding tab, edit or add the following parameters for the E-MapReduce Hook.
Tab
Parameter
Description
spark-defaults.conf
spark.sql.queryExecutionListeners
Captures SQL information from the Spark service for data lineage and access frequency.
-
To enable the E-MapReduce Hook, set this parameter to
com.aliyun.emr.meta.spark.listener.EMRQueryLogger. -
To disable the E-MapReduce Hook, leave the parameter value blank.
hive-site.xml
dlf.emrhook.webtracking
Specifies whether to report access frequency. Valid values:
-
true: enabled.
-
false: disabled.
NoteIf you disable the E-MapReduce Hook, the Data Overview page for data tables in the Data Lake Formation (DLF) console will no longer display data for File Visits within Last Day, File Visits within Last Seven Days, and File Visits within Last 30 Days.
-
-
Save the configuration.
-
On the Configure page, click Save.
-
In the dialog box that appears, enter a reason in the Execution Reason field and click Save.
-
-
-
Restart the Spark service.
-
On the Configure page, choose More > Restart.
-
In the dialog box that appears, enter a reason in the Execution Reason field and click OK.
-
In the Confirm dialog box, click OK.
-
-
View the data overview and data lineage.
-
View the Data Overview in Data Lake Formation (DLF). For more information, see Data overview of data tables.
-
View data lineage in DataWorks. For more information, see Data lineage analysis.
-
FAQ
Related documents
To configure the E-MapReduce Hook for the Hive service, see Use Hive extensions to record data lineage and access history.