This topic shows you how to log on to the E-MapReduce (EMR) console with your Alibaba Cloud account, create an EMR on ACK cluster, and run a job.
Notes
-
You are responsible for managing and configuring the runtime environment for your code.
-
In the examples in this topic, the JAR file is packaged directly into the image. If you use your own JAR file, you can upload it to Object Storage Service (OSS). For more information, see Simple upload.
Replace
local:///opt/spark/examples/spark-examples.jarin the command with the OSS path to your JAR file. The path format isoss://<yourBucketName>/<path>.jar.
Prerequisites
Before you create an EMR on ACK cluster, perform the following steps in the Container Service for Kubernetes (ACK) console:
-
Create a Kubernetes cluster. For more information, see Create an ACK dedicated cluster (new instances no longer available) or Create an ACK managed cluster.
-
Attach the AliyunOSSFullAccess and AliyunDLFFullAccess policies. For more information, see Grant OSS and DLF policies.
To store JAR files in Object Storage Service (OSS), you must first activate OSS. For more information, see Activate OSS.
Step 1: Grant role permissions
Before using EMR on ACK, grant the AliyunEMROnACKDefaultRole system default role to your Alibaba Cloud account in the EMR on ACK console. For more information, see Grant role permissions to an Alibaba Cloud account.
Step 2: Create a cluster
Create a Spark cluster on Kubernetes. For more information about how to create a cluster, see Create a cluster.
-
Log on to the EMR on ACK console.
-
On the EMR on ACK page, click Create Cluster.
-
On the E-MapReduce on ACK page, configure the cluster parameters.
Parameter
Example
Description
Region
China (Hangzhou)
The region where the cluster is deployed. You cannot change the region after the cluster is created.
Cluster Type
Spark
A general-purpose, distributed big data processing engine that provides capabilities such as ETL, offline batch processing, and data modeling.
ImportantAfter creating a Spark cluster, if you want to associate it with a Shuffle Service cluster, the product's major version must match that of the associated Shuffle Service cluster. For example, a Spark cluster of EMR-5.x-ack can be associated only with a Shuffle Service cluster of EMR-5.x-ack.
Product Version
EMR-5.6.0-ack
The latest software version is used by default.
Component Version
SPARK (3.2.1)
The components and their versions available for the selected cluster type.
ACK cluster
Emr-ack
Select an existing ACK cluster, or create one in the ACK console.
ImportantAn ACK cluster cannot be associated with multiple EMR on ACK clusters of the same type.
Click Configure Dedicated Nodes to configure dedicated nodes for EMR. This adds EMR-specific taints and labels to a node pool or specific nodes.
NoteWe recommend configuring dedicated nodes by using a node pool. If no node pool exists, create one. For more information, see Create a node pool and Node pools.
OSS bucket
oss-spark-test
Select an existing bucket, or create one in the Object Storage Service (OSS) console.
Cluster Name
Emr-Spark
The name of the cluster. It must be 1 to 64 characters long and can contain Chinese characters, letters, digits, hyphens (-), and underscores (_).
-
Click create.
The cluster is created when its status changes to Running.
Step 3: Submit a job
This section shows how to submit a Spark job by using a Custom Resource Definition (CRD). For more information about Spark, see the open source Quick Start documentation and select a suitable language and version.
For more information about submitting jobs, see the following topics:
Connect to an Alibaba Cloud Container Service for Kubernetes (ACK) cluster by using kubectl. For more information, see Connect to an ACK cluster using kubectl.
Create a job file named spark-pi.yaml. The following code shows the content in the file:
apiVersion: "sparkoperator.k8s.io/v1beta2" kind: SparkApplication metadata: name: spark-pi-simple spec: type: Scala sparkVersion: 3.2.1 mainClass: org.apache.spark.examples.SparkPi mainApplicationFile: "local:///opt/spark/examples/spark-examples.jar" arguments: - "1000" driver: cores: 1 coreLimit: 1000m memory: 4g executor: cores: 1 coreLimit: 1000m memory: 8g memoryOverhead: 1g instances: 1For information about the fields in the code, see spark-on-k8s-operator.
NoteYou can specify a custom file name. In this example, spark-pi.yaml is used.
In this example, Spark 3.2.1 for EMR V5.6.0 is used. If you use another version of Spark, configure the sparkVersion parameter based on your business requirements.
Run the following command to submit a job:
kubectl apply -f spark-pi.yaml --namespace <Namespace in which the cluster resides>Replace
<Namespace in which the cluster resides>with the namespace based on your business requirements. To view the namespace, log on to the EMR console and go to the Cluster Details tab.The following information is returned:
sparkapplication.sparkoperator.k8s.io/spark-pi-simple createdNotespark-pi-simpleis the name of the submitted Spark job.Optional. View the information about the submitted Spark job on the Job Details tab.
(Optional) Step 4: Release cluster
If you no longer need the cluster, release it to reduce costs.
-
On the EMR on ACK page, find the cluster and click Release in the Actions column.
-
In the Release Cluster dialog box, click OK.
Related documents
-
View a summary of your clusters: View cluster information.
-
View your cluster's jobs: View the job list.