All Products
Search
Document Center

E-MapReduce:Create a Data Science cluster

Last Updated:Sep 16, 2026

Create a Data Science cluster on EMR on ACK from the E-MapReduce (EMR) console.

Prerequisites

  • You have granted the AliyunOSSFullAccess and AliyunDLFFullAccess permissions. For more information, see Role authorization.

  • You have an existing Container Service for Kubernetes (ACK) cluster. For more information, see Create an ACK dedicated cluster (discontinued) or Create an ACK managed cluster.

  • Important

    Your ACK cluster must meet the following requirements:

    • Kubernetes version: 1.22 to 1.24.

    • vCPU: 16 or more.

    • Memory: 64 GiB or more.

    • Instance type:

      • Only general-purpose, compute-optimized, and memory-optimized instance types are supported.

      • Only the ecs.g5, ecs.g6, ecs.g7, and later instance families are supported.

  • You have created a node pool. For more information, see Create and manage node pools.

Precautions

You cannot deploy more than one Data Science cluster in the same ACK cluster.

Procedure

  1. Log on to EMR on ACK.

  2. On the EMR on ACK page, click Create Cluster.

  3. Configure the cluster parameters.

    Parameter

    Description

    Region

    The region in which to create the cluster. The region cannot be changed after creation.

    Cluster Type

    Data Science: Supports big data and AI scenarios, including offline big data ETL with Hive and Spark and model training with TensorFlow. You can select a CPU+GPU heterogeneous computing framework and use NVIDIA GPUs to accelerate certain deep learning algorithms.

    Product Version

    The latest software version is selected by default.

    Component Version

    Lists the components and versions for the selected cluster type.

    ACK cluster

    Select an existing ACK cluster, or create one in the Container Service for Kubernetes (ACK) console.

    Note

    Data Science clusters use the following namespaces: anonymous, cert-manager, fluid-system, ingress-nginx, istio-system, knative-serving, kubeflow, kubernetes-dashboard, and monitoring. If these namespaces already exist in your ACK cluster, creating the Data Science cluster overwrites them.

    Configure Dedicated Nodes

    Reserves a node pool or individual nodes for exclusive use by EMR by applying EMR-specific taints and labels.

    Cluster Name

    The name of the cluster. The name must be 1 to 64 characters long and can contain Chinese characters, letters, digits, hyphens (-), and underscores (_).

  4. Click create.

    The cluster is successfully created when its status changes to Running.