すべてのプロダクト
Search
ドキュメントセンター

E-MapReduce:Fusion エンジン

最終更新日:May 26, 2026

Fusion エンジンは、EMR Serverless Spark に組み込まれたベクトル化された SQL 実行エンジンです。TPC-DS ベンチマークテストでは、オープンソースの Spark よりも 3 倍優れたパフォーマンスを発揮します。Fusion エンジンはオープンソースの Spark と完全な互換性があるため、コードの変更は不要です。セッションの作成時に [Use Fusion Acceleration] スイッチをオンにすることで有効化できます。

Fusion エンジンは、Spark SQL と DataFrame のジョブを高速化します。ほとんどのオペレーター、式、データ型のパフォーマンスが向上します。

制限事項

Fusion エンジンは、以下の種類のジョブを高速化しません。

  • Resilient Distributed Dataset (RDD) ジョブ

  • ユーザー定義関数 (UDF) を使用するジョブ

サポートされているストレージフォーマット

  • Parquet

  • Paimon

  • ORC (部分的なサポート)

サポートされているオペレーター

タイプ

オペレーター

ソース

FileSourceScanExec, HiveTableScanExec, BatchScanExec, InMemoryTableScanExec

シンク

DataWritingCommandExec

共通操作

FilterExec, ProjectExec, SortExec, UnionExec

集約

HashAggregateExec

結合

BroadcastHashJoinExec, ShuffledHashJoinExec, SortMergeJoinExec, BroadcastNestedLoopJoinExec, CartesianProductExec

ウィンドウ

WindowExec, WindowTopK

Exchange

ShuffleExchangeExec, ReusedExchangeExec, BroadcastExchangeExec, CoalesceExec

Limit

GlobalLimitExec, LocalLimitExec, TakeOrderedAndProjectExec

サブクエリ

SubqueryBroadcastExec

その他

ExpandExec, GenerateExec

サポートされていないオペレーター

タイプ

オペレーター

集約

ObjectHashAggregateExec, SortAggregateExec

Exchange

CustomShuffleReaderExec

Pandas

AggregateInPandasExec, FlatMapGroupsInPandasExec, ArrowEvalPythonExec, MapInPandasExec, WindowInPandasExec

その他

CollectLimitExec, RangeExec, SampleExec

サポートされている式

タイプ

比較/論理

=, ==, !=, <>, <, <=, >, >=, <=>, is null, is not null, between, and, or, ||, !, negative, null if

算術

%, +, -, *, /, isnan, mod, negative, not, positive, abs, acos, acosh, asin, asinh, atan, atan2, atanh, cbrt, ceil, ceiling, cos, cosh, degrees, e, exp, floor, ln, log, log10, log2, pi, pmod, pow, power, radians, rand, random, rint, round, shiftleft, shiftright, sign, signum, sin, sqrt, tan, tanh

ビット単位

^, |, &, ~, bit_and, bit_count, bit_or, bit_xor, bit_length

条件

case, if, when

セット

in, find_in_set

文字列

ascii, char, chr, char_length, character_length, concat, instr, lcase, lower, length, locate, lpad, ltrim, overlay, replace, reverse, rtrim, split, split_part, substr, substring, trim, ucase, upper, like, regexp, regexp_extract, regexp_extract_all, regexp_like, regexp_replace, rlike

集約

aggregate, approx_count_distinct, avg, collect_list, collect_set, corr, count, covar_pop, covar_samp, first, first_value, kurtosis, last, last_value, max, max_by, mean, min, regr_avgx, regr_avgy, regr_count, regr_r2, regr_intercept, regr_slope, regr_sxy, regr_sxx, regr_syy, skewness, std, stddev, stddev_pop, stddev_samp, sum, var_pop, var_samp, variance

ウィンドウ

cume_dist, dense_rank, lag, lead, nth_value, ntile, percent_rank, rank, row_number

時間

add_months, current_date, current_timestamp, current_timezone, date, date_add, date_format, date_from_unix_date, date_sub, datediff, day, dayofmonth, dayofweek, dayofyear, from_unixtime, from_utc_timestamp, hour, last_day, make_date, minute, month, next_day, now, quarter, second, timestamp_micros, timestamp_millis, to_date, to_unix_timestamp, unix_seconds, unix_millis, unix_micros, weekday, weekofyear, year

JSON

get_json_object, json_array_length

配列

array, array_contains, array_distinct, array_except, array_intersect, array_join, array_max, array_min, array_position, array_remove, array_repeat, array_sort, arrays_overlap, arrays_zip, element_at, exists, filter, forall, flatten, shuffle, size, sort_array

Map

map, get_map_value, map_from_arrays, map_keys, map_values, map_zip_with, named_struct, struct, str_to_map

エンコーディング

crc32, hash, md5, sha1, sha2

その他

current_catalog, current_database, greatest, least, monotonically_increasing_id, nanvl, spark_partition_id, stack, uuid, rand

サポートされているデータ型

  • Byte、Short、Int、Long

  • ブール値

  • 文字列とバイナリ

  • 10 進数

  • Float と Double

  • Date とタイムスタンプ

サポートされていないデータ型

  • 構造体

  • 配列

  • Map

Fusion エンジンの有効化と使用

Fusion エンジンは、以下のいずれかの方法で有効にできます。

方法 1:セッション管理での有効化

SQL、ノートブック、または Spark Thrift Server のセッションを作成する際に、[Use Fusion Acceleration] オプションをオンにします。タスクを実行する際は、Fusion アクセラレーションが有効になっているセッションを選択して Fusion エンジンを使用します。

詳細については、「セッションの管理」をご参照ください。

方法 2:Spark パラメーターテンプレートの設定

EMR Serverless Spark コンソールで、[Configurations] > [Task Templates] に移動し、設定テンプレートに Fusion 関連のパラメーター (spark.emr.serverless.fusion または spark.emr.serverless.fusion.enabled) を追加します。Spark パラメーターの完全なリストについては、「カスタム Spark パラメーター」をご参照ください。

詳細については、「Spark 設定テンプレートの管理」をご参照ください。

方法 3:DataWorks での Spark パラメーター設定

Spark パラメーターは、ワークスペースレベルまたはノードレベルで設定できます。

グローバル設定

EMR タスクを実行する DataWorks モジュールに対して、ワークスペースレベルの Spark パラメーターを設定します。パラメーター spark.emr.serverless.fusion または spark.emr.serverless.fusion.enabled を追加し、値を true に設定します。Spark パラメーターの完全なリストについては、「カスタム Spark パラメーター」をご参照ください。

手順については、「グローバル Spark パラメーターの設定」をご参照ください。

ノードレベル設定

Data Studio の Spark ノードの場合、ノード編集ページの右側にあるスケジューリング設定で Spark パラメーターを設定します。パラメーター spark.emr.serverless.fusion または spark.emr.serverless.fusion.enabled を追加し、値を true に設定します。Spark パラメーターの完全なリストについては、「カスタム Spark パラメーター」をご参照ください。

手順については、「ノード Spark パラメーターの設定」をご参照ください。