All Products
Search
Document Center

E-MapReduce:Data lineage and access history with SparkSQL extensions

Last Updated:Jul 17, 2026

The E-MapReduce Hook collects data lineage and access frequency from SparkSQL jobs. You can use Data Lake Formation (DLF) to track table and partition access counts, or DataWorks to manage data lineage.

Prerequisites

A DataLake or custom cluster with the Spark service is required. For more information, see Create a cluster.

Limitations

  • EMR-HOOK does not support collecting SQL information from jobs in a gateway environment custom-deployed with EMR-CLI.

  • In EMR versions earlier than 5.16.0 and 3.50.0, the hive.exec.post.hooks (for Hive) and spark.sql.queryExecutionListeners (for Spark) parameter settings are not synchronized to gateway nodes. In EMR versions 5.16.0 and later and 3.50.0 and later, these settings are synchronized to gateway nodes. These versions also introduce the hive_aux_jars_path_gateway_only parameter, which lets you independently use custom extension JAR files on gateway nodes.

Usage notes

  • In E-MapReduce versions earlier than EMR-5.14.0 and EMR-3.48.0, the E-MapReduce Hook is enabled by default.

    You can manually enable the E-MapReduce Hook on custom clusters that run EMR-3.44, as it is disabled by default. For more information, see FAQ.

  • In E-MapReduce versions EMR-5.14.0 and later, and EMR-3.48.0 and later, the E-MapReduce Hook is disabled by default, and you must enable it manually.

Procedure

  1. Go to your cluster's Services page.

    1. Log on to the E-MapReduce console.

    2. In the top navigation bar, select a region and a resource group based on your requirements.

    3. On the EMR on ECS page, find the target cluster and click Services in the Services column.

  2. Configure the E-MapReduce Hook.

    1. On the Services page, find the Spark2 or Spark3 service and click Configure.

    2. On the Configure page, on the corresponding tab, edit or add the following parameters for the E-MapReduce Hook.

      Tab

      Parameter

      Description

      spark-defaults.conf

      spark.sql.queryExecutionListeners

      Captures SQL information from the Spark service for data lineage and access frequency.

      • To enable the E-MapReduce Hook, set this parameter to com.aliyun.emr.meta.spark.listener.EMRQueryLogger.

      • To disable the E-MapReduce Hook, leave the parameter value blank.

      hive-site.xml

      dlf.emrhook.webtracking

      Specifies whether to report access frequency. Valid values:

      • true: enabled.

      • false: disabled.

      Note

      If you disable the E-MapReduce Hook, the Data Overview page for data tables in the Data Lake Formation (DLF) console will no longer display data for File Visits within Last Day, File Visits within Last Seven Days, and File Visits within Last 30 Days.

    3. Save the configuration.

      1. On the Configure page, click Save.

      2. In the dialog box that appears, enter a reason in the Execution Reason field and click Save.

  3. Restart the Spark service.

    1. On the Configure page, choose More > Restart.

    2. In the dialog box that appears, enter a reason in the Execution Reason field and click OK.

    3. In the Confirm dialog box, click OK.

  4. View the data overview and data lineage.

FAQ

How do I enable the E-MapReduce Hook for a custom cluster that runs on EMR-3.44?

On the Configure tab for the Spark service, edit the following parameters and follow the prompts to apply the changes.

Tab

Parameter

Modification

spark-defaults.conf

spark.driver.extraClassPath

Append /opt/apps/EMRHOOK/emrhook-1.1.5/spark-hook-1.1.5-spark30.jar to the parameter value.

spark.executor.extraClassPath

Related documents

To configure the E-MapReduce Hook for the Hive service, see Use Hive extensions to record data lineage and access history.