All Products
Search
Document Center

E-MapReduce:EMR-3.x series releases

Last Updated:Sep 21, 2026

This topic provides the release notes for the E-MapReduce (EMR) 3.x series. For the supported components in each release version, see Release versions.

EMR-3.55.x

Release date

Version

Date

EMR-3.55.0

October 27, 2025

Updates

Service

Change

Ranger

  • Jindoauth Server now allows client users to use custom RAM roles to access OSS.

  • Fixed missing dependencies in the Ranger YARN plugin.

Paimon

Upgraded to 1-ali-16.3.

JindoCache

Upgraded to 6.10.1.

Component versions

DataLake cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

Hive

2.3.9

Spark 2

2.4.8

Spark 3

3.4.2

YARN

2.8.5

Trino

422

Delta Lake

3.0.0

Hudi

0.15.0

Iceberg

1.5.0

Flume

1.11.0

Kyuubi

1.9.2

Tez

0.10.2

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

Sqoop

1.4.7

DLF-Auth

2.0.2

Presto

0.283

ZooKeeper

3.8.4

Knox

1.5.0

Celeborn

0.5.2

JindoCache

6.10.1

Paimon

1-ali-16.3

OLAP cluster

Service

Version

StarRocks 2

2.5.22

StarRocks 3

3.2.11

Doris

2.1.4

ClickHouse

23.8.2.7

ZooKeeper

3.8.4

DataFlow cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

YARN

2.8.5

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

ZooKeeper

3.8.4

Knox

1.5.0

Flink

1.17.2

Paimon

1-ali-16.3

DataServing cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

ZooKeeper

3.8.4

Knox

1.5.0

HBase

1.7.1

JindoCache

6.10.1

Phoenix

4.16.1

Custom cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

Hive

2.3.9

Spark 2

2.4.8

Spark 3

3.4.2

YARN

2.8.5

Trino

422

Delta Lake

3.0.0

Hudi

0.15.0

Iceberg

1.5.0

Flume

1.11.0

Kyuubi

1.9.2

Tez

0.10.2

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

Sqoop

1.4.7

DLF-Auth

2.0.2

Presto

0.283

StarRocks 2

2.5.22

StarRocks 3

3.2.11

ZooKeeper

3.8.4

Knox

1.5.0

Celeborn

0.5.2

Flink

1.17.2

HBase

1.7.1

JindoCache

6.10.1

Paimon

1-ali-16.3

Phoenix

4.16.1

EMR-3.54.x

Release date

Version

Date

EMR-3.54.0

July 10, 2025

Updates

Service

Changes

Hive

Resolves known issues.

Tez

Incorporates community bug fixes for improved performance and stability.

Component versions

DataLake cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

Hive

2.3.9

Spark 2

2.4.8

Spark 3

3.4.2

YARN

2.8.5

Trino

422

Delta Lake

3.0.0

Hudi

0.15.0

Iceberg

1.5.0

Flume

1.11.0

Kyuubi

1.9.2

Tez

0.10.2

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

Sqoop

1.4.7

DLF-Auth

2.0.2

Presto

0.283

ZooKeeper

3.8.4

Knox

1.5.0

Celeborn

0.5.2

JindoCache

6.8.2

Paimon

1-ali-6.2

OLAP cluster

Service

Version

StarRocks 2

2.5.22

StarRocks 3

3.2.11

Doris

2.1.4

ClickHouse

23.8.2.7

ZooKeeper

3.8.4

DataFlow cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

YARN

2.8.5

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

ZooKeeper

3.8.4

Knox

1.5.0

Flink

1.17.2

Paimon

1-ali-6.2

DataServing cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

ZooKeeper

3.8.4

Knox

1.5.0

HBase

1.7.1

JindoCache

6.8.2

Phoenix

4.16.1

Custom cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

Hive

2.3.9

Spark 2

2.4.8

Spark 3

3.4.2

YARN

2.8.5

Trino

422

Delta Lake

3.0.0

Hudi

0.15.0

Iceberg

1.5.0

Flume

1.11.0

Kyuubi

1.9.2

Tez

0.10.2

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger plugin

1.0.0

Sqoop

1.4.7

DLF-Auth

2.0.2

Presto

0.283

StarRocks 2

2.5.22

StarRocks 3

3.2.11

ZooKeeper

3.8.4

Knox

1.5.0

Celeborn

0.5.2

Flink

1.17.2

HBase

1.7.1

JindoCache

6.8.2

Paimon

1-ali-6.2

Phoenix

4.16.1

EMR-3.53.x

Release date

Version

Date

EMR-3.53.0

April 24, 2025

Updates

Service

Change

Trino

Fixed an LDAP availability issue.

YARN

Fixed open-source defects (YARN-10213, YARN-6207, and YARN-9339).

StarRocks

Added support for storage-compute separated clusters.

JindoCache

Upgraded to 6.8.2.

EMRHOOK

Improved stability.

Release information

Data lake cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

Hive

2.3.9

Spark 2

2.4.8

Spark 3

3.4.2

YARN

2.8.5

Trino

422

Delta Lake

3.0.0

Hudi

0.15.0

Iceberg

1.5.0

Flume

1.11.0

Kyuubi

1.9.2

Tez

0.10.2

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger-plugin

1.0.0

Sqoop

1.4.7

DLF-Auth

2.0.2

Presto

0.283

ZooKeeper

3.8.4

Knox

1.5.0

Celeborn

0.5.2

JindoCache

6.8.2

Paimon

1-ali-6.2

OLAP cluster

Service

Version

StarRocks 2

2.5.22

StarRocks 3

3.2.11

Doris

2.1.4

ClickHouse

23.8.2.7

ZooKeeper

3.8.4

Data flow cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

YARN

2.8.5

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger-plugin

1.0.0

ZooKeeper

3.8.4

Knox

1.5.0

Flink

1.17.2

Paimon

1-ali-6.2

Data serving cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger-plugin

1.0.0

ZooKeeper

3.8.4

Knox

1.5.0

HBase

1.7.1

JindoCache

6.8.2

Phoenix

4.16.1

Custom cluster

Service

Version

Hadoop-Common

2.8.5

HDFS

2.8.5

OSS-HDFS

1.0.0

Hive

2.3.9

Spark 2

2.4.8

Spark 3

3.4.2

YARN

2.8.5

Trino

422

Delta Lake

3.0.0

Hudi

0.15.0

Iceberg

1.5.0

Flume

1.11.0

Kyuubi

1.9.2

Tez

0.10.2

OpenLDAP

2.4.46

Ranger

2.3.0

Ranger-plugin

1.0.0

Sqoop

1.4.7

DLF-Auth

2.0.2

Presto

0.283

StarRocks 2

2.5.22

StarRocks 3

3.2.11

ZooKeeper

3.8.4

Knox

1.5.0

Celeborn

0.5.2

Flink

1.17.2

HBase

1.7.1

JindoCache

6.8.2

Paimon

1-ali-6.2

Phoenix

4.16.1

EMR-3.52.x

Release date

Version

Date

EMR-3.52.1

December 18, 2024

EMR-3.52.0 (Not available for new purchases)

December 4, 2024

Updates

Service

Changes

Spark

  • Fixed configuration issues encountered during scale-out.

  • Fixed intermittent SASL connection failures in Kerberos clusters.

Hive

Fixed configuration issues encountered during scale-out.

Trino

Fixed an issue that prevented connections after enabling LDAP.

Presto

ZooKeeper

Added support for custom configuration.

Ranger

Replaced the Spark 3 Ranger plugin with the one from the open-source Kyuubi project.

Hudi

Upgraded to 0.15.0.

Celeborn

Upgraded to 0.5.2.

JindoCache

Upgraded to 6.5.3.

StarRocks 3

Upgraded to 3.2.11.

Kyuubi

Upgraded to 1.9.2.

StarRocks 2

Upgraded to 2.5.22.

Impala

The service is unavailable. You can use the recommended service as an alternative or manually install the corresponding service.

You can replace Impala with Presto, Trino, ClickHouse, or StarRocks.

Kudu

Kafka

Kafka-Manager

EMR-3.51.x

Release date

Version

Release date

EMR-3.51.4

December 18, 2024

EMR-3.51.3 (No longer available for new purchases)

November 29, 2024

EMR-3.51.2 (No longer available for new purchases)

August 29, 2024

EMR-3.51.1 (No longer available for new purchases)

June 21, 2024

EMR-3.51.0 (No longer available for new purchases)

April 23, 2024

Updates

EMR-3.51.4

Service

Change

JindoCache

Upgraded to 6.5.3.

StarRocks 2

Upgraded to 2.5.22.

StarRocks 3

Upgraded to 3.2.11.

EMR-3.51.3

Service

Description

JindoSDK

JindoSDK is updated to resolve the issue that causes deadlocks.

EMR-3.51.2

Service

Description

JindoCache

  • JindoCache is updated to 6.5.1.

  • The performance of reading data from and writing data to distributed hash tables is improved.

Spark

  • The issue that partition directories cannot be deleted is fixed.

  • The issue related to the Hive package dependency is fixed. This ensures that the connection between Spark and the Metastore client remains uninterrupted.

Trino

  • The issue that some modified configurations may be unexpectedly restored to original configurations during a scale-out is fixed.

  • Data in the OSS-HDFS service that is deployed in a high-security cluster can be queried.

  • The issue that exceptions occur on Trino after DLF-Auth is enabled is fixed.

Presto

Data in the OSS-HDFS service that is installed in a high-security cluster can be queried.

HDFS

The issue that the memory size of NameNodes and DataNodes cannot be modified is fixed.

HBase-HDFS

YARN

  • Multiple timeline events can be sent by the ResourceManager at a time, which improves the processing capability.

  • The logic issue in processing containers and resources of the ResourceManager is fixed.

ZooKeeper

  • The issue that the memory configuration of a node group cannot be modified is fixed.

  • The log configuration files can be reconstructed.

Impala

The issue that client configurations are unexpectedly modified during an auto scaling activity is fixed.

Ranger

The latest version of JindoSDK is supported, which effectively reduces the CPU load.

Knox

The following issue is fixed: The URL of Knox fails to be accessed when a cluster has only one Master Extend node group.

Kafka

The following issue is fixed: The EMR cluster in which Kafka Connect is deployed fails to be started.

StarRocks

The issue that added BE nodes are not displayed after a scale-out is fixed.

Doris

Doris is updated to 2.1.4.

Paimon

Paimon is updated to 0.9-ali-7.

EMR-HOOK

The lineage information of a MaxCompute table can be parsed.

EMR-3.51.1

Service

Change

Spark

Added support for deploying the Master-Extend node group.

Hive

Kyuubi

Paimon

Replaced the dependency on the VVR version of Flink with the community edition and added support for the DLF Catalog.

Knox

Packaged with JDK 8.

Flink

Restored the Data Lake Formation (DLF) configurations and dependencies removed in EMR-3.51.0.

EMR-3.51.0

Service

Change

Spark

Upgraded Spark 3 to 3.4.2.

Celeborn

Upgraded to 0.4.0.

Doris

Upgraded to 2.1.0.

StarRocks

  • Upgraded StarRocks 2 to 2.5.18.

  • Upgraded StarRocks 3 to 3.2.4.

Delta Lake

Upgraded to 3.0.0.

Iceberg

Upgraded to 1.5.0.

ZooKeeper

Upgraded to 3.8.4.

JindoCache

Upgraded to 6.2.5.

Flink

Upgraded to 1.17.2.

EMR-3.50.x

Release date

Version

Date

EMR-3.50.0

February 19, 2024

Updates

Service

Changes

Hudi

Upgraded to version 0.14.0.

Flume

Upgraded to version 1.11.0.

Kyuubi

Upgraded to version 1.7.3.

Impala

Upgraded to version 4.3.0.

Celeborn

Upgraded to version 0.3.2.

JindoCache

Upgraded to version 6.2.0.

Paimon

Upgraded to version 0.7-ali-1.

Kafka

  • Upgraded to version 3.6.1.

  • Fixed a SASL authentication vulnerability in Kafka Connect.

Spark

Fixed a vulnerability in Commons Text.

StarRocks

  • Upgraded StarRocks 2 to version 2.5.13.

  • Upgraded StarRocks 3 to version 3.1.5.

Ranger

  • Fixed a vulnerability in Commons Text.

  • Fixed a path-matching permission bypass vulnerability in Spring Security.

  • Fixed a forward/include authentication bypass vulnerability in Spring Security.

  • Fixed a special pattern-matching authentication bypass vulnerability in Spring Framework.

  • The Ranger LDAP user synchronization period is now configurable.

EMR-3.49.x

Release date

Version

Date

EMR-3.49.1

November 16, 2023

EMR-3.49.0 (Not available for purchase)

October 27, 2023

Updates

Component

Change

JindoCache

Added the JindoCache component, version 6.1.1.

JindoData

JindoData is unavailable. You can use JindoCache for data caching and DLF-Auth for authentication.

Spark

Removed jdo-related configuration from hive-site.xml.

HBase

Added a configuration option to select the HBase Thrift Server version (v1 or v2).

StarRocks

Upgraded StarRocks to 2.5.10.

Doris

Upgraded Doris to 1.2.7.

Celeborn

Upgraded Celeborn to 0.3.1.

Paimon

Upgraded Paimon to 0.6-ali-2.

ClickHouse

Upgraded ClickHouse to 23.8.2.7.

EMR-3.48.x

Release date

Version

Date

EMR-3.48.2

August 17, 2023

Updates

Service

Description

Trino

  • Fixed an issue where the Paimon connector could not query HDFS tables.

  • Fixed an issue where worker node metrics could not be read.

Presto

  • Upgraded to version 0.283.

  • Fixed an issue where worker node metrics could not be read.

ClickHouse

The default user is now granted all permissions.

StarRocks

  • Renamed the existing StarRocks service to StarRocks2.

  • Added StarRocks3 (version 3.1.2), which uses compute-storage integration by default. Compute-storage separation is not supported.

Celeborn

Upgraded to version 0.3.0.

EMR-3.47.x

Release date

Version

Release date

EMR-3.47.0

August 3, 2023

Updates

Service

Description

Hudi

Upgraded to version 0.13.1.

Paimon

Upgraded to version 0.5-ali-1.

StarRocks

Upgraded to version 2.5.8.

JindoData

Upgraded to version 4.6.11.

Trino

  • Upgraded to version 422.

  • Added support for querying merge on read (MOR) tables with the Hudi connector.

  • Improved error messages for dynamic UDF loading.

EMR-3.46.x

Release date

Version

Date

EMR-3.46.1

July 13, 2023

EMR-3.46.0 (Not available for purchase)

June 1, 2023

Updates

EMR-3.46.1

Service

Description

Spark

  • By default, OSS-HDFS is used to store data of Spark History Server.

  • OSS or OSS-HDFS is used to store data of Spark3 Native Engine.

Hive

By default, OSS-HDFS is used to store data in Hive warehouse files.

OSS-HDFS

The OSS-HDFS service is added.

YARN

By default, OSS-HDFS is used to store data.

HBase

  • By default, OSS-HDFS is used to store HBase data in the HFile format.

  • OSS-HDFS is used to store write-ahead logging (WAL) logs of HBase.

EMR-3.46.0

Service

Description

Kyuubi

Upgraded to version 1.7.1.

Celeborn

Upgraded to version 0.2.2.

Paimon

  • Flink Table Store is renamed to Paimon.

  • Upgraded to version 0.4-ali-1.

StarRocks

Upgraded to version 2.5.5.

Doris

Upgraded to version 1.2.4.

ClickHouse

Upgraded to version 22.8.17.17.

Trino

Includes a simple event listener by default to retrieve audit logs.

Phoenix

Adds support for Hive on Phoenix.

EMR-3.45.x

Release dates

Version

Date

EMR-3.45.1

April 3, 2023

EMR-3.45.0 (Not available for new purchases)

February 28, 2023

Updates

EMR-3.45.1

Service

Description

ClickHouse

ClickHouse is updated to 22.8.14.53.

Trino

The odps.properties connector is added. This allows you to query MaxCompute data.

JindoData

JindoData is updated to 4.6.5.

JindoSDK

JindoSDK is updated to 4.6.5.

Flink Table Store

Flink Table Store is updated to 0.3-ali-2.

YARN

The Node Labels feature is supported.

EMR-3.45.0

Service

Description

Iceberg

Upgraded to version 1.1.0.

Hudi

  • Upgraded to version 0.12.2.

  • Introduced support for Change Data Capture (CDC).

Kudu

Upgraded to version 1.16.0.

ClickHouse

  • Upgraded to version 22.3.8.39.

  • When installing the ClickHouse service, you must also select ZooKeeper.

Celeborn

  • Remote Shuffle Service (RSS) is renamed to Celeborn.

  • Celeborn is now version 0.2.0.

Presto

The Presto service is now available. It uses the Facebook PrestoDB 0.278.3 community kernel and defaults to port 8889 for HTTP and port 7779 for HTTPS.

Delta Lake

Upgraded to version 2.2.0.

StarRocks

Upgraded to version 2.4.3.

Doris

Upgraded to version 1.2.1.

Kafka Manager

Upgraded to version 3.0.0.6.

Impala

The service is retired.

OpenLDAP

Upgraded to version 2.4.46.

Kyuubi

Upgraded to version 1.6.1.

Ranger

Upgraded to version 2.3.0.

HBase

  • Now supports ThriftServer2.

  • The default value of the hbase.block.data.cachecompressed parameter is now true.

Flink Table Store

Introduced the Flink Table Store service, based on community version 0.3.

JindoData

Upgraded to version 4.6.4.

EMR-3.44.x

Release date

EMR-3.44.0 was released on December 1, 2022.

Changes

Service

Changes

Iceberg

Upgraded to version 0.14.1.

Flink

Upgraded to Flink 1.15-vvr-6.0.2, which is based on the Apache Flink 1.15 major version.

Kafka

  • Added support for user logon authentication and authorization using LDAP.

  • Added support for user group authorization.

Trino

  • EMR Presto is renamed to Trino, its official community name.

  • Added support for Ranger and DLF-Auth.

  • Fixed an issue where connections to worker nodes failed after LDAP was enabled.

JindoSDK

Upgraded to version 4.6.2.

JindoData

Upgraded to version 4.6.2.

HBase

  • Added support for Ranger.

  • Fixed an issue that prevented OSS-HDFS from being selected as the storage mode when adding the service.

YARN

The Access Control List (ACL) feature is now enabled by default in security mode.

StarRocks

Upgraded to version 2.3.4.

Doris

Upgraded to version 1.1.5.

Hudi

Added support for configuring hudi-defaults.conf in the console.

Ranger

Added integration with Trino, YARN, HBase, and Kafka.

DLF-Auth

  • Upgraded to version 2.0.2.

  • Added support for Trino and Impala.

OpenLDAP

Integrated with the Nslcd component.

Kudu

The Kudu TServer can no longer be installed on a task node group.

Spark

Upgraded to version 3.3.1.

Tez

Upgraded to version 0.10.2.

Kyuubi

Upgraded to version 1.6.0.

EMR-3.43.x

Release dates

Version

Date

EMR-3.43.1

November 8, 2022

EMR-3.43.0 (Not available for new purchases)

October 14, 2022

Updates

EMR-3.43.1

Service

Change

Kerberos

You can now connect EMR to an external KDC.

Kafka

Added a startup command option to customize service startup parameters.

JindoData

  • Upgraded to 4.6.0.

  • You can now rewrite the OSS-HDFS access path.

Flink

Upgraded to 1.13_vvr_4.0.15.

RSS

Upgraded to 0.1.4.

EMR-3.43.0

Service

Change

Spark

  • Upgraded to 3.3.

  • Now supports Kerberos authentication.

Hudi

  • Upgraded to 0.12.0.

  • Added support for Spark 3.3.

  • You can now host metadata on a cloud metastore and enable the acceleration feature. For more information, see Hudi metastore usage guide.

Flink

  • Now supports Kerberos authentication.

  • Now supports automatic connections to Data Lake Formation (DLF).

Iceberg

  • Upgraded to 0.14.0.

  • Added support for Spark 3.3.

  • Now supports Kerberos authentication.

JindoData

  • Upgraded to 4.5.1.

  • You can now access Alibaba Cloud resources without a plaintext AccessKey.

Hadoop-Common and HDFS

  • Now supports Kerberos authentication.

  • Fixed vulnerability CVE-2022-25168.

Knox

Integrated with Ranger. You can access the Ranger UI from the Access Links and Ports tab.

HBase

  • Upgraded to 1.7.1.

  • Now supports Kerberos authentication.

  • Now supports node group configuration.

RSS

  • Upgraded to 0.1.2.

  • Now supports Kerberos authentication.

Doris

  • Upgraded to 1.1.2.

  • Now supports Kerberos authentication.

StarRocks

  • Upgraded to 2.2.6.

  • Now supports Kerberos authentication.

Kafka

  • Upgraded to 2.13_3.2.1.

  • Now supports Kerberos authentication.

Delta Lake

  • Upgraded to 2.1.0.

  • Added support for Spark 3.3.

  • Now supports Kerberos authentication.

Kudu

The Kudu component, version 1.14.0, is now available.

Impala

  • You can now create views by using Data Lake Formation (DLF).

  • Now supports Kerberos authentication.

YARN, Impala, Ranger, Hive, Kyuubi, Tez, Kafka, ZooKeeper, DLF-Auth, Phoenix, Sqoop, and Presto

Now supports Kerberos authentication.

EMR-3.42.x

Release date

EMR-3.42.0 was released on August 5, 2022.

Changes

Service

Changes

Hive

Added support for one-click LDAP integration.

Presto

  • Upgraded to community version 389.

    Uses the community's standalone connectors for Delta Lake and Hudi.

    • The Delta Lake connector in this version does not support the Time Travel or Z-Order features.

    • The Hudi connector does not support querying MOR tables.

  • Added support for one-click LDAP integration.

Delta Lake

  • Now integrates with Data Lake Formation to automatically manage lake tables.

  • Now supports Ranger authentication.

  • Fixed an issue that prevented the collection of statistics for TIMESTAMP fields.

  • The optimize and vacuum commands now return metrics.

Hudi

Upgraded to version 0.11.1.

Hadoop Common

Added the Hadoop Common component to resolve configuration conflicts among HDFS, YARN, and JindoSDK.

YARN

Enhanced auto scaling.

Ranger

  • Now supports both Spark 2 and Spark 3.

  • Now supports one-click LDAP integration with Ranger UserSync.

Kafka

Cruise Control now automatically creates the required topics on startup.

HBase

Added the HBase component, version 1.4.9.

Phoenix

Added the Phoenix component, version 4.14.1.

Doris

Upgraded to version 1.1.1.

StarRocks

Upgraded to version 2.2.3.

ClickHouse

Fixed an out-of-memory (OOM) error when reading large files from Object Storage Service.

EMR-3.40.x

Release date

EMR-3.40.0 April 21, 2022

Updates

Component

Description

JindoData

Added as a new component in version 4.3.0.

JindoSDK

Upgraded to 4.3.0.

Spark

Upgraded to 3.2.1.

Hive

  • Fixed an issue causing duplicate commits in Tez when speculative execution was enabled.

  • Fixed an issue where a UDF had to be reloaded before being called.

Presto

Fixed an issue where the Presto service failed to start when added to an initialized Hadoop cluster.

Delta Lake

Fixed a compatibility issue with Streaming SQL.

Hudi

Upgraded to 0.10.1.

Iceberg

Upgraded to 0.13.1.

YARN

  • Added a configuration option to restrict Application Masters (AMs) to core nodes.

  • Fixed an issue where taihaodoctor was missing from the mapreduce.map.java.opts configuration.

ZooKeeper

Optimized JVM parameters.

Flink

Adapted for JindoSDK 4.3.0.

Impala

Flume

Druid

Sqoop

Upgraded the PostgreSQL version.

Zeppelin

Fixed an issue where the JDBC interpreter failed to start.

Ranger

The Ranger 1.2.0 Spark plugin now supports Hudi.

Oozie

Upgraded Log4j to 2.17.2.

HBase

Fixed an issue where the RegionServer for HBase 1.4.9 failed to start.

DLF-Auth

Upgraded to 2.0.0.

EMR-3.39.x

Release dates

Version

Date

EMR-3.39.2

March 25, 2022

EMR-3.39.1 (Not available for purchase)

February 15, 2022

Updates

EMR-3.39.2

Note

This version is only available for OLAP clusters and Dataflow clusters in the new EMR console.

Service

Updates

Flink

  • Enhanced the APM dashboard and added new metrics, such as sourceIdleTime.

  • Added support for CloudMonitor alerts.

Kafka

  • Added support for SSL and SASL configuration.

  • Changed the default values of some parameters.

ClickHouse

Changed the default values of some parameters.

EMR-3.39.1

Service

Updates

SmartData

Retired the components.

BIGBOOT

RSS

  • Upgraded EMR Remote Shuffle Service (ESS) to RSS. For more information, see RSS.

  • Enhanced the functionality and stability of the service.

JindoSDK

  • Upgraded the architecture to JindoData.

  • This is the first release to integrate JindoSDK 4.0 in EMR, with support for OSS and the OSS-HDFS service.

Spark

  • Optimized Hive on Spark.

  • Adapted Spark for JindoSDK.

Tez

Adapted Tez for JindoSDK.

Hive

Adapted Hive for JindoSDK.

Presto

  • Added support for dynamic UDF loading.

  • Added support for time travel queries on Delta Lake tables by using the FOR ... AS OF syntax.

  • Added a standalone Delta Lake catalog with default Delta connector settings. This catalog also supports Z-order data skipping optimization.

  • Fixed an issue where the Hudi connector could not query Hudi MOR tables. The Hive connector still does not support querying Hudi MOR tables.

  • Adapted Presto for JindoSDK.

Delta Lake

  • Metadata management

    • Delta Lake now uses the built-in Spark catalog to synchronize metadata and partition information instead of the Hive CLI API.

    • Delta Lake now automatically reports table statistics (data profiling) to the metastore.

  • SQL

    • Added support for time travel syntax.

    • Added support for DROP PARTITION syntax.

    • Added support for ADD COLUMN operations at specified positions (FIRST and AFTER).

  • Enhanced table management capabilities

    • Added a feature to dynamically adjust file sizes based on the table size. This feature is enabled by default.

    • Enabled automatic vacuum by default and added support for concurrent vacuum operations.

    • Optimized the auto-compaction logic. This feature is disabled by default.

    • Added Z-order syntax and improved Z-order performance.

Hudi

Upgraded to 0.10.0.

HDFS

Adapted HDFS for JindoSDK.

YARN

Adapted YARN for JindoSDK.

Flume

Adapted Flume for JindoSDK.

Flink

  • By default, the lib directory of Flink is uploaded to the HDFS cluster, which allows you to use the directory with the yarn.provided.lib.dirs parameter.

  • Adapted Flink for JindoSDK.

Impala

Adapted Impala for JindoSDK.

Ranger

  • Fixed an issue where Spark History Server failed to start.

  • Adapted Ranger for JindoSDK.

HBase

  • Fixed default parameter issues.

  • Fixed an issue with the date format in GC logs.

  • Fixed a restart issue that occurred when using an IP address for a RegionServer.

Druid

Adapted Druid for JindoSDK.

ClickHouse

Optimized the logic for stopping the ClickHouse component.

Iceberg

  • Upgraded to 0.13.0.

  • Hid default configuration items to improve usability.

DLF-Auth

Fixed an issue where Spark History Server failed to start.

StarRocks

This service is now available in the new EMR console.

Version 2.0.1 is available.

E-MapReduce (EMR) 3.38.x

Release date

Version

Release date

EMR-3.38.3

December 2021

EMR-3.38.2 (Not available for purchase)

December 2021

EMR-3.38.1 (Not available for purchase)

November 2021

EMR-3.38.0 (Not available for purchase)

October 2021

Updates

EMR-3.38.3

Fixed the Log4j security vulnerability in all related components. For details, see Security bulletin | Apache Log4j2 remote code execution vulnerability.

Service

Updates

Presto

  • Fixed an issue where Presto in a high-availability cluster failed to query Hudi tables.

  • Fixed the Log4j vulnerability in the Elasticsearch connector.

Data Lake Formation (DLF) metastore

  • Disabled metastore logging by default.

  • Fixed an error caused by an excessively long gettablestats URI in the metastore.

Delta Lake

Fixed an issue where schema changes were not synchronized to the metastore.

Flink

  • Upgraded VVR to 4.0.11. This version includes the following features:

    • Introduced commercial Flink CDC features:

      • Added support for Schema Evolution.

      • Added support for Flink SQL semantics for entire database synchronization.

    • Added support for using GeminiStateBackend to store state in Object Storage Service (OSS).

  • Added an Enterprise Edition Hudi connector with built-in Data Lake Formation (DLF) for metadata management.

Sqoop

Fixed precision loss for the DECIMAL type when Sqoop imported data to HCatalog tables.

EMR-3.38.2

Service

Updates

SmartData

  • Upgraded SmartData to 3.8.0. For details, see Introduction to SmartData 3.8.x.

  • Added support for Kerberos-based authentication and Ranger-based authorization for Object Storage Service (OSS).

EMR-3.38.1

Service

Updates

SmartData

Upgraded SmartData to 3.7.3. For details, see Introduction to SmartData 3.7.x.

Spark

  • Removed the invalid Log4j MetricsAppender configuration.

  • Fixed a null pointer exception during SparkContext startup.

Presto

  • Fixed an issue where Presto in a high-availability Hadoop cluster required a host configuration to query Hive tables.

  • Fixed an issue where Presto could not start with the default configuration on low-memory nodes.

  • Fixed an issue where changes to the worker-jvm configuration did not take effect.

  • Added support for Ranger.

Impala

Fixed an issue where a no such method error occurred when you queried a DLF metadata table.

Ranger

  • Added support for Presto.

  • Fixed permission issues for Ranger Spark inserts into ORC and PARQUET tables.

  • Fixed an issue where Ranger Hive role permissions did not take effect after Kerberos was enabled.

Data Lake Formation (DLF) authentication

  • Upgraded DLF authentication to 1.0.1.

  • Added support for using Data Lake Formation (DLF) to control Presto permissions.

  • Fixed a caching issue that affects RAM users.

EMR-3.38.0

Service

Updates

SmartData

Upgraded SmartData to 3.7.2. For details, see Introduction to SmartData 3.7.x.

Spark

  • Upgraded Spark to 2.4.8.

  • Supports both Spark 2.4.8 and Spark 3.1.2.

    Note

    Spark 3 does not support Delta Lake or Remote Shuffle Service.

  • In Spark 3.x, SparkSQL optimizes DISTINCT computations. The optimization is triggered when an aggregation operator contains multiple count(distinct case ... when ...) expressions.

  • Fixed an array index out of bounds issue in Adaptive Query Execution (AQE) when stats were missing.

  • Fixed issues with AQE and caching in specific scenarios.

Hive

Upgraded Hive to 2.3.9.

Presto

  • Added support for standalone Presto clusters.

  • Upgraded Presto to community version 358.

    Important

    This version does not support Ranger.

  • Added default support for connectors such as Hudi and MySQL, and updated the default configurations.

  • Enabled auto scaling for Presto clusters.

  • Added support for data lake analysis.

Delta Lake

  • Unified Delta Lake connectors for Hive 2 and Hive 3.

  • Fixed an error in Delta Lake connectors when querying multi-level partitioned tables.

Hudi

  • Upgraded Hudi to 0.9.0.

  • Fixed a compatibility issue with sql.extension between Delta Lake and Hudi.

HDFS

Increased the default reserved space for NameNode to adaptively enter Safe mode when disk space runs low.

Flink

  • Upgraded Flink to 1.13-vvr-4.0.10, which corresponds to community Flink 1.13.1.

  • Added commercial Flink connectors, such as a Hologres connector.

  • Added metric reporters to integrate with APM dashboards.

  • For the Kafka connector, added a SchemaRegistry-based Kafka catalog. You can read data from and write data to existing Kafka topics without using DDL.

Storm

Retired the component.

Zeppelin

Upgraded Zeppelin to community version 0.10.0.

Ranger

This version of Ranger does not support permission control for Presto community version 358.

Hue

  • Fixed an issue where YARN Job Browser could not display or terminate jobs in some cases.

  • Enabled YARN Job Browser in the default configuration.

  • Enabled the Presto protocol in the default configuration.

Druid

Fixed an issue where leftover PID files after a power outage could prevent nodes from restarting.

ClickHouse

  • Updated the default configuration.

  • Added support for cluster scale-out.

  • Supports the MetaChecker feature.

  • Added support for reading data using the Object Storage Service (OSS) table engine and table function.

  • Added support for table-level custom ZooKeeper addresses.

Iceberg

Added the component. The version is 0.12.0-1.0.1.

Knox

Fixed an issue where the first access to a Spark task failed.

Data Lake Formation (DLF) authentication

Added the component.

You can now use Data Lake Formation (DLF) to control Hive and Spark permissions. The component version is 1.0.0.

Auto Scaling

Upgraded Auto Scaling to 1.2.0.

EMR-3.37.x

Release dates

Version

Release date

EMR-3.37.1

September 2021

EMR-3.37.0 (Not available for purchase)

August 2021

Updates

EMR-3.37.1

Service

Description

SmartData

Upgraded SmartData to version 3.7.1.

Hue

Fixed an issue where Impala could not be used in a high-security cluster.

Kudu

Added support for Kerberos.

EMR-3.37.0

Service

Description

SmartData

Upgraded SmartData to version 3.7.0.

Spark

Fixed a compatibility issue with Delta Lake.

Delta Lake

  • Upgraded Delta connectors to support table creation and queries using the StorageHandler syntax.

  • Fixed an issue that occurred during an INSERT OVERWRITE operation on a partitioned table.

  • Fixed an issue where the OPTIMIZE command wrote virtual fields to files in the G-SCD scenario.

YARN

  • Added appId, CPU, and memory resource usage information to the container REST APIs on nodes.

  • Fixed an issue where Application Master (AM) logs could not be viewed on nodes released by auto scaling.

  • Added support for cleaning up decommissioned nodes after auto scaling.

  • Enhanced the graceful decommission logic for auto scaling. The system now marks a node as decommissioned only after its NodeManager process ends.

ZooKeeper

Upgraded to community version 3.6.3.

Flink

  • Added the SmartData component.

  • Fixed an issue that prevented password-free access to OSS when using SSH to connect to a Dataflow Flink cluster and submit a job.

Impala

Fixed a directory listing loop that occurred when directly deleting an OSS partition directory.

Hue

Fixed a UI display issue when using Hue with Oozie.

Kudu

Upgraded to community version 1.14.0.

ClickHouse

Updated the default configuration.

EMR-3.36.x

Release date

EMR-3.36.1 was released on July 16, 2021.

Updates

Service

Change

SmartData

Upgraded SmartData to version 3.6.1.

For details, see Introduction to SmartData 3.6.x.

Hive

  • Upgraded Hive to version 2.3.8.

  • Fixed an issue where the show create table command returned an incorrect result when using DLF (Data Lake Formation) metadata.

  • Optimized the default Hive parameters to improve job performance.

  • Capitalized parameter names on the hive-env tab of the Hive service's Configure page in the EMR console for improved usability.

  • The error message that is reported because of the incompatibility between the file system and Hive metastore when you write data to a Hive table is optimized.

HDFS

Added support for the Zstandard (ZSTD) compression format.

Flink

Upgraded Flink to version 1.12-vvr-3.0.2.

Note

Removed Flink from Hadoop clusters.

Hudi

  • Upgraded Hudi to version 0.8.0.

  • Added support for Spark SQL integration.

Spark

  • Optimized parameter names on the spark-defaults tab of the Spark service's Configure page in the EMR console.

  • Improved log output performance.

  • Added support for the Zstandard (ZSTD) compression format.

Impala

Fixed a core dump error that occurred when using HDFS.

Tez

Optimized the default Tez parameters to improve job performance.

Knox

  • Added support for Kudu.

  • Added support for Impala.

  • Added support for HBase.

Phoenix

Fixed the "JDBC Driver not found" error when accessing Phoenix tables with Hive or Spark SQL.

ClickHouse

Added Application Performance Management (APM) for monitoring and alerting.

EMR-3.35.x

Release date

EMR-3.35.0 was released on April 21, 2021.

Updates

Service

Change

SmartData

Upgraded to version 3.5.0.

For more information, see Introduction to SmartData 3.5.x.

Spark

  • Fixed an issue where Adaptive Execution did not take effect in some scenarios.

  • Fixed an issue where the behavior of statistical aggregate functions was inconsistent with that of Hive.

  • Fixed an issue where data of the CHAR type was read incorrectly from Hive ORC tables.

HDFS

HDFS now supports the SM4 encryption algorithm.

Hue

Upgraded Hue to version 4.9.0.

Alluxio

Upgraded Alluxio to version 2.5.0.

Druid

  • Upgraded Druid to version 0.20.1.

  • Enhanced security.

Livy

Upgraded Livy to version 0.7.1.

E-MapReduce 3.34.x

Release date

E-MapReduce 3.34.0 was released on March 15, 2021.

Updates

Service

Updates

SmartData

Upgraded to version 3.4.0.

For details, see SmartData 3.4.x Release Notes.

Spark

  • Optimized some default configurations.
  • Performance optimization: Window TopK pushdown is supported.
  • Enhanced the compatibility of Hive when it reads and writes CSV or JSON tables.
  • The ANALYZE statement supports omitting all column names of a table.
  • Supports enabling or disabling the LDAP feature with a single click.
  • Improved the usability of the Spark Beeline tool.

Hive

  • Optimized default configurations.

  • Enhanced performance by improving the cost-based optimizer (CBO).

  • Added a one-click option to enable or disable the LDAP feature.

  • Upgraded Calcite to version 1.12.0.

  • Added the hive.security.authorization.sqlstd.confwhitelist.append parameter.

Presto

Added a one-click option to enable or disable the LDAP feature.

YARN

Fixed a high-risk vulnerability that allowed unauthorized access to the Hadoop web UI. This issue required specifying user.name=name in the URL when accessing the YARN web UI through an SSH tunnel.

ZooKeeper

Upgraded to version 3.6.2.

Flink

Updated the config.sh file during initialization to fix an issue with HADOOP_CLASSPATH.

Impala

  • Upgraded to version 3.4.0.

  • Upgraded Shiro to version 1.7.0.

  • Added support for DLF metadata.

  • Enabled querying data in the Delta format.

  • Added a one-click option to enable or disable the LDAP feature.

Tez

Optimized default configurations.

HAS

Fixed an issue where the admin.keytab file failed to re-initialize after an error during the HAS installation.

Ranger

  • The issue caused by filter pushdown in Spark is fixed.

  • The issue that prevents Presto from being enabled after you disable Presto in Ranger is fixed.

  • LDAP authentication can be enabled or disabled with a click.

Knox

Fixed a connection issue with Knox for Druid 0.20.0.

Hue

Added a one-click option to enable or disable the LDAP feature.

Hudi

  • Supports the SQL on Hudi feature.
  • Fixed an issue where query results were inaccurate for some data.
  • Supports partition pruning when Spark queries Copy On Write tables in Hudi.
  • Supports the bucket index mechanism to improve write performance.

Delta Lake

  • Fixed an issue where metadata could not be synchronized to the Hive Metastore for existing Delta tables.
  • Fixed an issue where the Merge command could not parse *.
  • Fixed an issue where an error was reported when Parquet-formatted data was converted to a Delta table and the table metadata was created.
  • Fixed an issue where the Optimize command ran abnormally when no files were pending compaction.
  • The Merge syntax supports using a subquery as the source.
  • Introduced a caching mechanism to improve query efficiency when Presto queries Delta tables.
  • Supports querying Delta tables by using Impala.

Superset

  • The issue that prevents the admin user from logging on to the web UI is fixed.

  • Datasets are compatible with Druid clusters.

  • Spark SQL datasets are no longer supported.

Sqoop

Enabled importing Parquet files to OSS.

Alluxio

Upgraded to version 2.4.1.

Phoenix

Hive on Phoenix now supports field settings.

Pig

Removed.

EMR-3.33.x

Release date

EMR-3.33.0 was released on January 15, 2021.

Updates

Service

Description

SmartData

Upgraded to version 3.2.0.

For more information, see Introduction to SmartData 3.2.x.

Spark

  • Upgraded to version 2.4.7.

  • Upgraded jQuery to version 3.5.1.

  • Added support for automatically updating Hive table and partition sizes.

  • Added support for exporting Spark metadata and job run information to DataWorks.

Hive

  • Upgraded to version 2.3.7.

  • HCatalog now supports Data Lake Formation.

  • Added support for exporting Hive metadata and job run information to DataWorks.

metastore

  • Added the Hive statistics feature.

  • HCatalog now supports Data Lake Formation.

  • Optimized the method for obtaining STS tokens.

HDFS

Upgraded jQuery to version 3.5.1.

YARN

  • Upgraded jQuery to version 3.5.1.

  • Adjusted Fair Scheduler configurations.

  • Optimized the Timeline Server.

Zeppelin

Upgraded to version 0.9.0.

Ranger

  • Added audit log configurations for Hive.

  • Added configurations for Log4j Audit.

OpenLDAP

  • Added an audit feature.

  • Enabled SSL port 10636 by default.

  • Added one-click enablement for Presto.

Knox

  • Fixed vulnerabilities in the Spring framework.

  • Fixed an issue on the Executors page of the Spark UI.

  • Fixed an issue on the job status page of the Oozie UI.

Hue

Added support for Presto.

Druid

Upgraded to version 0.20.0.

EMR Hook

  • Introduced as a new service.

  • hive-hook: Added support for exporting Hive metadata and job run information to DataWorks.

  • spark-hook: Added support for exporting Spark metadata and job run information to DataWorks.

EMR-3.32.x

Release date

EMR-3.32.0 was released on November 23, 2020.

Updates

Service

Change

SmartData

Upgraded to 3.1.0.

See Introduction to SmartData 3.1.x.

Alluxio

  • Added support for Alluxio 2.4.0.

  • The default configuration scales with cluster node size.

  • By default, EMR clusters use HDFS as the UnderFS, which works out-of-the-box.

  • Enhanced the Alluxio OSS UnderFS to support new features such as OSS versioning.

  • Now compatible with engines such as Hadoop, Hive, Spark, and Presto.

Hudi

Added support for Hudi 0.6.0.

Spark

JindoTable now supports enabling or disabling data collection.

Hive

  • Fixed a connection pool leak in HiveServer.

  • JindoTable now supports enabling or disabling data collection.

  • Optimize the performance of ADD COLUMN.

  • Fixed incorrect data reads from Hudi tables.

  • The default configuration scales with cluster node size.

HDFS

Increased the snapshot limit.

YARN

The default configuration scales with cluster node size.

Tez

The default configuration scales with cluster node size.

Sqoop

Fixed an issue with Avro file imports.

EMR-3.30.x

Release date

EMR-3.30.0 was released on October 26, 2020.

Updates

Service

Updates

SmartData

Upgraded to 3.0.0.

For more information, see Introduction to SmartData 3.0.x.

Spark

  • Added support for metadata from Alibaba Cloud Data Lake Formation (DLF).

  • Upgraded HAS dependencies to 2.0.1.

  • Fixed an issue with backticks in Streaming SQL.

  • Removed the Delta JAR package. Delta is now deployed separately.

  • Log paths now write to HDFS.

Hive

  • Added support for metadata from Alibaba Cloud Data Lake Formation (DLF).

  • Fixed an issue where a DUMMY file was written when reading an empty directory from a Delta table.

  • Upgraded HAS dependencies to 2.0.1.

Presto

  • Added support for metadata from Alibaba Cloud Data Lake Formation (DLF).

  • Removed limitations on reading Delta tables.

  • Fixed missing JVM configurations in high security mode.

  • Upgraded HAS dependencies to 2.0.1.

HDFS

  • Added support for hot-swap disk mode.

  • Upgraded HAS dependencies to 2.0.1.

YARN

  • Fixed issues with YARN RMZKStateStore.

  • Added support for SNAPPY files from Log Service (SLS).

  • Modified the directory configuration for MapReduce local mode to fix a directory permission check issue.

  • Added support for hot-swap disk mode.

  • Logs are now written to HDFS.

  • Upgraded HAS dependencies to 2.0.1.

ZooKeeper

  • Added support for binding service ports to an internal IP address.

  • Upgraded HAS dependencies to 2.0.1.

Flink-Vvp

  • Upgraded to 1.11-2.2.2.

  • Added support for SQL and Autopilot.

Note

Flink-Vvp is supported only on Dataflow clusters. Hadoop clusters do not currently support Flink-Vvp.

Flink

  • Added support for writing data to OSS in cache mode. This feature, combined with Flink checkpoints and a replayable source, provides EXACTLY_ONCE semantics.

  • Synced with Apache Flink 1.11.1. SQL now supports multi-way output (MULTI INSERT).

  • Upgraded HAS dependencies to 2.0.1.

Impala

  • Added support for customizing catalogd.flgs, impalad.flgs, and statestored.flgs.

  • Upgraded Shiro to 1.6.0.

  • Upgraded HAS dependencies to 2.0.1.

Tez

  • Optimized the default AM memory parameters.

  • Upgraded HAS dependencies to 2.0.1.

HAS

Upgraded HAS dependencies to 2.0.1.

Storm

Zeppelin

Ranger

OpenLDAP

Oozie

Knox

Kafka

HUE

HBase

Druid

EMR-3.29.x

Release date

EMR-3.29.0 was released on July 29, 2020.

Updates

Service

Description

Bigboot

  • Upgraded to version 2.7.301.

  • Jindo DistCp now supports writing data to OSS in archive or infrequent access mode.

  • Enhanced Fuse to support multiple namespaces.

  • Improved metadata caching in cache mode.

Spark

  • Upgraded Spark to version 2.4.5.2.0.

  • Added support for third-party metastores.

  • Added the datalake metastore-client.

Hive

  • Upgraded Hive to version 2.3.5.6.0.

  • Added support for third-party metastores.

  • Added the datalake metastore-client.

Presto

Upgraded to version 338.

Ranger

  • Upgraded the software package to version 1.2.0-1.5.0.

  • Added support for Presto 338.

  • Added a description field to the configuration file.

HDFS

The reserved space for DataNodes is now adaptively configured.

Knox

Added support for Impala, later versions of Flink, and Platform for AI (PAI).

Druid

Upgraded to version 0.18.1.

SmartData

Upgraded to version 2.7.301.

EMR-3.28.x

Release date

EMR-3.28.0 was released on June 12, 2020.

New features

Service

Changes

Bigboot

  • Released JindoTable, which supports statistics collection based on table or partition popularity.

  • Added support for complete storage policies in block storage mode and for tiered storage policies, including infrequent access and archive.

  • Added the Jindo DistCp data migration tool.

  • Improved Jindo Fuse and fixed related issues.

  • Improved the integration of the JFS scheme with the Hive engine and Jindo JobCommitter in cache mode.

  • In block storage mode, you can now configure a weight to read data directly from OSS. This helps distribute and reduce the overhead of reading from the local cache.

  • Decoupled JindoFS into separate software modules: Bigboot (management layer), SmartData (distributed service), and JindoFS SDK. Each module can be upgraded and maintained independently.

Updates

Service

Changes

Flink

Upgraded open source Flink to Ververica Platform. This platform is a customized version of Flink 1.10 that provides additional features, such as the proprietary Gemini storage engine.

Bigboot

Upgraded to version 2.7.0.

Delta

  • Upgraded to version 0.6.0.

  • Decoupled the Delta code from the Spark code.

Spark

  • Upgraded to version 2.4.5.

  • Added compatibility with DataFactory's streaming-sql scripts.

  • Added support for Delta 0.6.0.

Hive

Added support for Delta 0.6.0.

Ranger

  • Added support for custom deployment of HDFS, Hive, and Spark.

  • You can now configure ranger-admin-site.xml and ranger-ugsync-site.xml in the console.

HDFS

If no DataNode is available when you write data to HDFS, HDFS now displays an exception message (HDFS-9023) identifying the unavailable DataNode.

Hue

  • Added support for installing the Hue component on gateway clusters.

  • Added support for deploying multiple Hue instances on a single node.

DataFactory

Added support for Delta 0.6.0.

Druid

Upgraded to version 0.18.0.

Knox

  • Upgraded to version 1.1.0-1.0.7.

  • Added Knox support for the HBase UI.

EMR-3.27.x

Release dates

Version

Release date

EMR-3.27.0

April 29, 2020

EMR-3.27.1 (Not available for purchase)

May 8, 2020

EMR-3.27.2 (Not available for purchase)

May 20, 2020

New features

Feature

Description

Custom service deployment

You can now customize the deployment of services on a master node. The following services are supported:

  • Hadoop

  • Spark

  • Hive

  • ZooKeeper

  • Presto

Graceful shutdown for auto scaling

After you enable graceful shutdown, EMR releases a node only after it completes its running tasks within a specified time.

Updates

Service

Description

Spark

  • Added support for date-type partition fields in CUBE.

  • Increased the stack depth for spark-submit.

Delta

  • Enhanced DDL statements, including CREATE, SHOW, and DESCRIBE.

  • Added support for the OPTIMIZE statement with Z-order.

Knox

  • Knox is now compatible with the Druid UI.

  • Added support for multi-master deployment.

Hive

  • Added support for the magic committer for HCatalog tables.

  • Removed outdated default configurations.

Bigboot

  • Upgraded to version 2.6.3.

  • Added support for multi-master deployment.

SmartData

  • Upgraded to version 2.6.3.

  • Added support for multi-master deployment.

Ranger

  • Added support for the Solr service.

  • Added support for PrestoSQL version 311.

Tez

Added support for setting the scratch directory (scratchdir) on OSS.

Presto

Upgraded to version 331.

Druid

Upgraded to version 0.17.1.

Superset

Upgraded to version 0.35.2.

Sqoop

  • Upgraded the MySQL JDBC JAR package to version 5.1.48.

  • The MySQL direct export mode now supports custom encoding with the --mysql-charset parameter.

EMR-3.26.x

Release date

Version

Release date

EMR-3.26.3 (not available for purchase)

April 16, 2020

Updates

Service

Changes

Bigboot

  • Upgraded to 2.6.3.

  • Added support for Tablestore metadata and namespace high availability.

SmartData

Hive

Enabled support for a direct committer for HCatalog tables.

YARN

Configured JindoOssCommitter as the default committer.

HDFS

Updated JindoFS-related configurations.

Spark

Configured JindoOssCommitter as the default committer.

EMR-3.25.x

Release date

EMR-3.25.0 was released on January 13, 2020.

New features

Ranger service: Added support for Presto operations.

Updates

Service

Changes

Ranger

  • Initialized the RangerAdmin database in high-availability clusters.

  • Fixed a security issue in the RangerUserSync startup script.

Spark

  • Added support for configuring Delta-related parameters, such as spark.sql.extensions, in the console.

  • Added support for reading Delta tables by using Hive without setting the input format.

  • Added support for ALTER TABLE SET TBLPROPERTIES and UNSET TBLPROPERTIES statements.

Delta

Hive

Fixed an issue where MapReduce tasks failed in automatic local mode.

Presto

  • Upgraded to version 310.

  • Upgraded joda-time to version 2.10.5.

Tez

  • Upgraded to version 0.9.2.

  • Fixed an issue where application progress was displayed incorrectly on the Tez UI.

  • Fixed an issue that prevented viewing the application history on the Tez UI.

Impala

Fixed an issue where Impala could not access LZO tables.

HDFS

Removed JAR packages related to mongo-hadoop.

ZooKeeper

Upgraded to version 3.5.6.

YARN

You can now add the yarn.resourcemanager.system-metrics-publisher.enabled=true configuration item on the yarn-site.xml tab to support the Tez UI.

Bigboot

  • Upgraded to version 2.2.3.

  • Added support for the rename operation in OSS cache mode.

SmartData

Knox

Upgraded dependency packages.

Oozie

Upgraded dependency packages.

EMR-3.24.x

Release date

EMR-3.24.0 was released on November 18, 2019.

New features

Service

Changes

Delta

  • Added support for SQL statements, including ALTER, CONVERT, CREATE, CTAS, DELETE, DESC, INSERT, MERGE, OPTIMIZE, UPDATE, and VACUUM.

  • Built in and optimized the OPTIMIZE command.

  • Added support for the Hive connector.

  • Added support for other open source features.

Grafana

Added the Grafana component, version 6.4.2, for Flink independent clusters.

Prometheus

Added the Prometheus component, version 2.13.0, for Flink independent clusters.

Alertmanager

Added the Alertmanager component, version 0.19.0, for Flink independent clusters.

TensorFlow on Spark

  • Added support for deploying the TensorFlow framework on Spark. This deep integration enables an end-to-end process from data preprocessing to deep learning training, featuring optimized task scheduling and data exchange.

  • Added support for streaming tasks.

Updates

Service

Changes

SmartData

  • Optimized the usage modes of JindoFS. The block mode is unchanged. The cache mode is enhanced for compatibility with existing OSS file system usage. It now supports data caching and metadata caching, both of which are disabled by default and can be configured separately.

  • Optimized read and write performance in block mode and cache mode.

  • Optimized disk cleanup. The system now measures hot data cached on local disks more accurately and clears it more promptly to ensure that disk usage does not exceed the configured quota.

  • Enhanced support for gateway clusters, enabling the use of both block and cache modes.

  • Added support for a deployment mode that separates a single storage cluster from multiple computing clusters.

Spark

  • Added support for Delta-related parameters.

  • Added support for configuring the Ranger Spark plug-in.

  • Upgraded JindoCube to version 0.3.0.

Hive

  • Added a SQL compatibility check feature.

  • Released the Hive 2.3.5 and Hadoop 2.8.5 combination.

  • The content of the hiveserver2-site.xml file is no longer synchronized to the hive-site.xml file in the spark-conf directory when the component is restarted.

  • Added support for using the MSCK command to add incremental directories.

  • Fixed a bug that occurred when Hive reused a Tez container.

  • Added support for using the MSCK command to optimize column directories.

Bigboot

Upgraded to version 2.2.1, which fixes an issue where native code could not be executed on specific server models.

Ranger

  • Refactored the deployment mode of the Spark plug-in.

  • Fixed a bug where the emr-header-2 node in a high-availability cluster could not obtain a keytab.

Kudu

Fixed the startup logic.

ZooKeeper

Added configuration for four-letter word commands, which are enabled by default.

HDFS

Adapted HDFS for JindoFS.

YARN

  • Changed the default value of yarn.scheduler.capacity.node-locality-delay to -1.

  • Adapted YARN for JindoFS.

Has

Integrated with OpenLDAP as the backend.

OpenLDAP

Adapted OpenLDAP for Has.

Presto

Upgraded to version 0.228.

Kafka

Introduced a feature that automatically isolates bad disks in the d1 instance family.

Druid

Upgraded to version 0.16.0.

Flume

Upgraded to version 1.9.0.

Flink

  • Upgraded to version 1.9.1.

  • Added support for Flink independent clusters (available via whitelist).

EMR-3.23.x

Release date

EMR-3.23.0 was released on September 18, 2019.

Updates

Service

Description

Druid

  • Upgraded to 0.15.1.

  • Added the router component.

  • Upgraded fastjson.

Spark

  • Fixed a class loader issue in the spark thrift server.

  • Improved the stability of transaction-related code in Spark.

  • Fixed read/write issues with the ORC format after upgrading the built-in Hive to version 2.3.

  • Added support for the MERGE INTO syntax.

  • Added support for the SCAN and STREAM syntax.

  • The Structured Streaming Kafka sink now supports Exactly-Once Semantics (EOS).

  • Upgraded Delta to 0.4.0.

Hive

  • Removed legacy Hive hooks.

  • Optimized for data skew in queries that contain multiple COUNT(DISTINCT) fields.

  • Fixed a data loss issue when joining tables with different bucket versions.

Flink

Upgraded to 1.8.2.

Bigboot

  • Updated the small file tool.

  • Fixed an issue in the OSS JAR that caused a non-daemon thread to remain active.

Kafka

  • Added Deployment Set awareness.

  • Removed the fastjson dependency.

HDFS

  • Optimized the deployment logic of the SmartData OSS JAR.

  • Updated the SmartData OSS JAR.

Flume

Upgraded fastjson.

Tensorflow on Spark

Added Tensorflow on Spark.

Has

Upgraded fastjson.

Livy

Upgraded fastjson.

EMR-3.22.x

Release date

EMR-3.22.0 was released on July 28, 2019.

New features

Component

Changes

Kudu

  • Added the Kudu component. Kudu fills a gap in the Hadoop ecosystem. It offers fast data insertion, modification, and random access similar to HBase, and supports large-scale data analysis and queries like HDFS or Parquet.

    • Provides C++ and Java APIs for custom development.

    • Provides integration with Impala, Spark, and Hive Metastore.

  • This version is based on open source Apache Kudu 1.10.0.

OpenLDAP

  • Added the OpenLDAP component to replace ApacheDS. ApacheDS is now retired.

  • Supports high availability.

Updates

Component

Description

JindoFileSystem

  • Multiple storage modes

    • Block mode: In this mode, data is stored as blocks in the backend OSS, and the local Namespace service maintains the metadata. This mode offers superior performance for metadata and data operations. Block mode supports various storage policies, including WARM (local and OSS replicas), COLD (OSS-only replica), HOT (multiple local replicas and one OSS replica), TEMP (local-only replica), and ALL_HDD (multiple local replicas). The default policy is WARM. You can set different storage policies for directories based on your application scenarios.

    • Cache mode: This mode is primarily compatible with existing OSS storage methods. In cache mode, files are stored as objects in OSS. Based on access patterns, data and metadata are cached locally to improve access performance. This mode provides various metadata synchronization policies for different scenarios.

  • External client support

    • The client SDK allows you to access the JindoFS file system from outside an E-MapReduce cluster. You can use the client to access the namespace in block mode. However, external clients cannot use the data cache within the E-MapReduce cluster, which results in lower performance than access from within the cluster.

    • Cache mode retains the original OSS storage semantics and accelerates data caching within an E-MapReduce cluster by using JindoFS. Therefore, you can directly access data from outside the E-MapReduce cluster by using an OSS client, such as the OSS SDK or the OssFileSystem of E-MapReduce.

  • Ecosystem component support

    • JindoFS now supports various computing engines in E-MapReduce, such as Spark, Flink, Hive, MapReduce, Impala, and Presto.

    • In storage-computing separation scenarios, JindoFS can store job logs, such as YARN container logs and Spark event logs.

    • JindoFS can be used as the backend storage for HBase HFiles to extend the storage capacity of HBase.

OssFileSystem

  • Added an automatic mechanism to detect bad disks and prevent resulting cache write failures to OSS.

  • Added the required configurations for OssFileSystem.

Bigboot

  • Upgraded to version 2.0.0.

  • Includes major updates such as support for multiple namespaces, storing local data blocks as large files, multi-mode storage, and external clients.

  • Fixed an issue where the Bigboot monitor displayed an incorrect status during a server restart.

  • Added a service spec for the Kudu component.

  • Added correctness validation for service specs.

Hadoop

  • HDFS

    • Adapted for HDFS Federation. You can now create an HDFS Federation cluster by using custom configurations and API operations. This avoids reformatting the NameNode when you create a Federation cluster.

    • Optimized the bad disk detection logic. For local disk scenarios, you can run the dfsadmin command to trigger a DataNode block report.

  • YARN

    Fixed an issue where the MapReduce JobHistory job list was not updated when MR job container logs were stored in JindoFS or OSS.

Spark

  • Relational Cache

    Added support for Relational Cache to accelerate user queries through pre-computation. You can create a relational cache to pre-compute data. When a query is executed, Spark Optimizer automatically finds a suitable cache, rewrites the SQL execution plan, and continues computation based on the cached data to improve query speed. This feature is suitable for scenarios such as reporting, dashboards, data synchronization, and multidimensional analysis.

    • Supports CACHE, UNCACHE, ALTER, and SHOW operations by using DDL statements. Cached data supports all data sources and data formats in Spark.

    • Supports automatic cache data updates and manual updates by using the REFRESH command. Incremental, partition-based updates are also supported.

    • Supports execution plan optimization based on Relational Cache.

  • Streaming SQL

    • Standardized the parameter configuration of Stream Query Writer.

    • Optimized the schema compatibility check for Kafka data tables.

    • A schema is automatically created in Schema Registry if a schema does not exist for a Kafka data table.

    • Improved the log messages for Kafka schema incompatibility.

    • Fixed an issue where column names had to be explicitly specified when query results are written to a Kafka table.

    • Removed the restriction that streaming SQL queries only support Kafka and Loghub as data input sources.

  • Delta

    Added support for Delta. You can use Spark to create a Delta data source to support scenarios such as streaming data writes, transactional reads and writes, data validation, and time travel. For more information, see the Delta documentation.

    • Supports reading data from and writing data to Delta by using the DataFrame API.

    • Supports using Delta as a source or sink for reading or writing data by using the Structured Streaming API.

    • Supports update, delete, merge, vacuum, and optimize operations by using the Delta API.

    • Supports creating Delta-based tables, importing data to Delta, and reading Delta tables by using SQL.

  • Other updates

    • Added support for constraints, including primary key and foreign key.

    • Resolved servlet and other JAR package conflicts.

Flink

Reverted the log4j logging configuration.

Kafka

  • Reverted the log4j logging configuration.

  • Upgraded fastjson.

Zeppelin

Upgraded the commons-lang3 dependency package to version 3.7 to fix an issue where PySpark could not write to OSS. For more information, see Spark 2.4 incompatibility with commons-lang3 in Zeppelin.

Ranger

Added support for the SHOW GRANTS statement.

Analytics-Zoo

Fixed a NumPy installation error.

Impala

Impala is now compatible with Apache Kudu 1.10.0.

Presto

Upgraded to version 0.221.

ZooKeeper

Upgraded to version 3.5.5.

Versions earlier than EMR-3.22.x

EMR-3.1.1

  • Upgraded the operating system to CentOS 7.2.

  • Upgraded Spark to version 2.1.1.

  • Upgraded EMR-Core to version 1.2.6.

  • Resolved an issue with AccessKey-free operations on OSS.

EMR-3.0.2

  • Upgraded EMR-Core to version 1.2.5.

  • Expanded AccessKey-free access to OSS to include more regions.

  • Adjusted the AccessKey replacement policy for RAM roles.

  • Fixed several Hive and Hadoop issues.

EMR-3.0.1

  • Introduced interactive mode and unified table management. Hive metadata is now stored in a unified external database, allowing multiple clusters to share the same metadata.

  • Upgraded EMR-Core to version 1.2.4 to improve the read and write performance of OSS.

  • Upgraded Spark to version 2.0.2.

Note

EMR-3.0.1 is fully compatible with EMR-3.0.0.

EMR-3.0.0

Initial release of EMR.