All Products
Search
Document Center

Object Storage Service:Use Spark on an EMR cluster to process data stored in OSS-HDFS

Last Updated:Sep 24, 2026

This topic describes how to use Spark on an E-MapReduce (EMR) cluster to process data stored in OSS-HDFS.

Prerequisites

Procedure

  1. Log on to the E-MapReduce console. In the left-side navigation pane, click EMR on ECS and create an EMR cluster.

    When you create the EMR cluster, make sure that Product Version is EMR-3.46.2 or later, or EMR-5.12.2 or later, and Root Storage Directory of Cluster is set to an OSS-HDFS-enabled bucket. Use the defaults for other parameters. For details, see Create a cluster.

  2. Run the following command on the terminal to start Spark Shell:

    spark-shell
  3. Use Spark to access OSS-HDFS.

    1. Create a table.

      spark.sql("CREATE TABLE test_oss (`c1` string) OPTIONS (PATH 'oss://examplebucket.cn-hangzhou.oss-dls.aliyuncs.com/dir')")
    2. Insert data into the table.

      spark.sql("INSERT INTO TABLE test_oss SELECT 'testdata' AS c1")
    3. Query data in the table.

      spark.sql("SELECT c1 FROM test_oss")