This topic describes how to use Spark on an E-MapReduce (EMR) cluster to process data stored in OSS-HDFS.
Prerequisites
-
A cluster of EMR V3.42.0 or later, or EMR V5.8.0 or later is created. For more information, see Create a cluster.
-
OSS-HDFS is enabled for a bucket and access permissions on OSS-HDFS are granted. For more information about how to enable OSS-HDFS, see Enable OSS-HDFS and grant access permissions.
Procedure
-
Log on to the E-MapReduce console. In the left-side navigation pane, click EMR on ECS and create an EMR cluster.
When you create the EMR cluster, make sure that Product Version is EMR-3.46.2 or later, or EMR-5.12.2 or later, and Root Storage Directory of Cluster is set to an OSS-HDFS-enabled bucket. Use the defaults for other parameters. For details, see Create a cluster.
-
Run the following command on the terminal to start Spark Shell:
spark-shell -
Use Spark to access OSS-HDFS.
-
Create a table.
spark.sql("CREATE TABLE test_oss (`c1` string) OPTIONS (PATH 'oss://examplebucket.cn-hangzhou.oss-dls.aliyuncs.com/dir')") -
Insert data into the table.
spark.sql("INSERT INTO TABLE test_oss SELECT 'testdata' AS c1") -
Query data in the table.
spark.sql("SELECT c1 FROM test_oss")
-