All Products
Search
Document Center

E-MapReduce:What is E-MapReduce

Last Updated:Aug 27, 2026

E-MapReduce (EMR) is an open source big data platform on Alibaba Cloud built on Apache Hadoop and Apache Spark. Use EMR to process and analyze data with the Hadoop and Spark ecosystems, and transfer data between Alibaba Cloud services such as Object Storage Service (OSS) and ApsaraDB RDS.

Deployment options

EMR offers the following deployment options:

Type

Description

EMR on ECS

Deploys open source Hadoop ecosystem components on Elastic Compute Service (ECS) instances. Manage clusters and services from the EMR console.

What is EMR on ECS?

EMR on ACK

Deploy a Container Service for Kubernetes (ACK) cluster, then run EMR big data components in containers on ACK resources. What is EMR on ACK?

EMR Serverless Spark

EMR Serverless Spark is a fully managed lakehouse engine for data and AI workloads. It is 100% compatible with open-source Spark -- run existing jobs with spark-submit and spark-sql without code changes.

What is EMR Serverless Spark?

Important

EMR is built on community open-source services such as Apache Hadoop, Spark, Flink, and StarRocks. Responsibility for O&M differs by deployment option:

  • EMR on ECS and EMR on ACK are semi-managed. Cluster resources belong to your Alibaba Cloud account, and you are responsible for the routine O&M of the open-source services, including capacity planning, parameter tuning, and troubleshooting. Plan for in-house big data O&M expertise to keep your workloads running.

  • EMR Serverless Spark and EMR Serverless StarRocks are fully managed. Alibaba Cloud operates the underlying resources and engine services, so you only develop jobs and analyze data.

For the full support boundary, see Scope and methods of technical support.

Benefits

EMR on ECS

Enterprise-grade open source big data services. Deploy Hadoop, Spark, Flink, Kafka, and HBase clusters quickly.

  • 100% community open source components, adapted and optimized for performance beyond stock releases.

  • Time-based elastic scaling with Spot Instance support for cost reduction.

  • Decoupled compute and storage for elastic resource utilization.

  • Create and scale clusters in minutes with no manual setup.

EMR on ACK

  • Cost-effective: No separate ACK cluster required.

  • Simplified O&M: Unified cluster management across big data and online services.

  • Flexible infrastructure: Supports both ECS and ACK resource models with seamless switching.

  • Deep integration: Cloud-native data lake architecture with ACK-based compute for unlimited scaling.

EMR Serverless Spark

  • Cloud-native high-speed compute engine

    • Built-in Fusion Engine (Spark Native Engine): Delivers 300% higher performance than the open source version and significantly accelerates big data computing jobs. A vectorized engine and batch data processing technologies optimize computing efficiency and reduce memory usage, which greatly improves overall performance.

    • Built-in Celeborn (Remote Shuffle Service): Supports petabyte-scale shuffle data processing and greatly improves the stability and performance of large shuffle jobs. Compute nodes do not require large-capacity disks. The dynamic resource scaling capability of Spark is fully utilized to reduce storage costs and lower the total cost of computing resources by up to 30%.

  • Open data lake architecture

    • On-demand elastic scaling: Supports a storage-compute separation architecture. Computing resources can be scaled within seconds, with a minimum granularity of one core. Resources are metered at a fine-grained job or queue level. Storage uses the pay-as-you-go billing method to avoid resource waste and greatly reduce enterprise operating costs.

    • Seamless migration and compatibility: Integrates with OSS-HDFS, a cloud storage service that is fully compatible with HDFS, to support smooth migration of your workloads to the cloud. DLF provides comprehensive metadata interconnection for the lakehouse to ensure data access consistency and complete permission management, which helps you build a modern data lakehouse architecture with ease.

  • One-stop developer experience

    • End-to-end development support: Provides a one-stop development experience that covers job development, debugging, release, and scheduling to meet the high standards of enterprise-grade development and release. The built-in version management feature records the complete release history and supports comparison of source code and configuration differences to ensure that changes are traceable.

    • Efficient collaboration and stability: Development and production environments are strictly isolated to ensure business stability and to support efficient team collaboration and stable delivery.

  • Serverless resource platform

    • Ready to use: You can quickly start job development without manual management or complex infrastructure setup.

    • Second-level elasticity: Resources are dynamically provisioned to start pods based on the resource requirements of Spark jobs, and are released immediately after computing is complete. You are charged only for the resources that you actually use, which further reduces the total computing cost.

    • Cost estimation: Provides job-level resource metering and cost estimation to help enterprises achieve fine-grained operations.