Before creating a project space for data development, you must configure a compute engine for your Dataphin instance. After the compute engine is configured, you can add the corresponding compute source to a project space to provide compute and storage resources.
Permissions
Only a super administrator or a system administrator can configure compute engines.
Billing
To configure a real-time compute engine, you must first purchase and enable the Real-time R&D module.
The Agile R&D Edition does not support real-time compute engines.
Limitations
-
After configuring a compute engine for a business tenant, reconfiguring the metastore compute engine type may cause incorrect metadata processing for that tenant. We recommend that you contact the Dataphin operations team for confirmation before you modify the metastore compute engine type.
-
When you modify the offline compute engine settings, the system automatically updates the compute source configuration. To ensure efficiency, the system does not verify the connectivity of the compute source during this process. Ensure the configuration is accurate to prevent task failures. After the modification is complete, we recommend that you manually test the connectivity of the compute source.
-
After you modify the compute settings, the new configuration takes effect on the compute source within 30 seconds. Before the synchronization is complete, you may see inconsistencies when viewing the compute source configuration, and SQL execution may still use the previous settings.
Supported compute engines
In single-engine mode, configure a compute engine for your Dataphin instance by specifying its cluster address. Once configured, you can create compute sources based on this cluster. In single-tenant multi-engine mode, you can only use a compute engine by associating it with a cluster. For more information, see Create a cluster in single-tenant multi-engine mode. Dataphin supports the following compute engines:
If no offline compute source exists, you can change both the compute engine type and its configuration. If one already exists, you can only modify the configuration, not the engine type.
If the tenant's metastore compute engine is already initialized, you can only select compute engines supported by that metastore.
|
Compute engine |
Description |
References |
|
Offline compute engine |
||
|
MaxCompute |
An Alibaba Cloud-native big data computing platform that provides efficient, stable storage and computing for massive datasets. |
Set MaxCompute as the compute engine for a Dataphin instance |
|
AnalyticDB for PostgreSQL |
A cloud-hosted, petabyte-scale, high-concurrency real-time data warehouse for online analytical processing (OLAP), with seamless scaling for massive data computation. |
Set AnalyticDB for PostgreSQL as the compute engine for a Dataphin instance |
|
E-MapReduce 3.x Hadoop and E-MapReduce 5.x Hadoop |
A Hadoop cluster based on Alibaba Cloud E-MapReduce (EMR), running on ECS instances. |
|
|
CDH 5.x Hadoop CDH 6.x Hadoop |
A widely used distributed system framework. Its core components, HDFS and MapReduce, provide massive data storage and computation. |
|
|
A widely used distributed system framework. Its core components, HDFS and MapReduce, provide massive data storage and computation. |
||
|
Cloudera Data Platform 7.x |
Cloudera Data Platform (CDP) combines the best of Cloudera CDH and Hortonworks HDP, created after their merger. |
|
|
Huawei FusionInsight 8.x Hadoop |
An enterprise-grade big data platform from Huawei for data storage, query, and analysis, with enhancements based on open source Apache software. |
|
|
AsiaInfo DP 5.3 Hadoop |
An integrated platform for big data production and operations, built on an open source ecosystem and carrier-grade technology. |
|
|
Transwarp ArgoDB |
A distributed analytical database from Transwarp. Note
Transwarp ArgoDB is not supported in the Intelligent R&D Edition. |
Set TDH or ArgoDB as the compute engine for a Dataphin instance |
|
Transwarp Data Hub (TDH) 6.x |
A big data platform from Transwarp. |
|
|
StarRocks |
A high-performance analytical data warehouse that uses vectorization, a Massively Parallel Processing (MPP) architecture, a cost-based optimizer (CBO), smart materialized views, and a real-time updatable columnar storage engine for multi-dimensional, real-time, high-concurrency data analysis. |
Use StarRocks as the metastore compute engine for initialization |
|
Lindorm (compute engine) |
An Alibaba Cloud-native multi-model database whose compute engine mode supports offline big data applications. |
Set Lindorm (compute engine) as the compute engine for a Dataphin instance |
|
GaussDB (DWS) |
GaussDB (DWS) is a distributed relational database developed by Huawei. It is based on PostgreSQL and is compatible with Oracle, MySQL, and Teradata syntax. |
Set GaussDB (DWS) as the compute engine for a Dataphin instance |
|
Databricks |
A unified data analytics platform based on Apache Spark. It provides managed Spark clusters, an interactive notebook environment, and seamless integration with cloud storage for efficient data processing and large-scale analytics. |
Set Databricks as the compute engine for a Dataphin instance |
|
Amazon EMR |
A managed cluster platform for running big data frameworks such as Hive and Spark. |
Set Amazon EMR as the compute engine for a Dataphin instance |
|
SelectDB |
SelectDB’s commercial distribution of Apache Doris. |
Set SelectDB or Doris as the compute engine for a Dataphin instance |
|
Doris |
Apache Doris is a high-performance, real-time analytical database based on an MPP architecture. |
|
|
EMR Serverless Spark |
A high-performance lakehouse product for Data + AI. |
Set EMR Serverless Spark as the compute engine for a Dataphin instance |
|
OushuDB |
A cloud-native distributed analytical database developed by Oushu Technology. |
Important
The OushuDB compute engine can only be used through a cluster in multi-engine mode. |
|
Real-time compute engine |
||
|
Realtime Compute for Apache Flink |
An Alibaba Cloud service based on Apache Flink that supports both real-time and offline batch processing with high throughput and low latency. |
After you enable the Real-time R&D module for a tenant, the system recommends a configuration based on the selected offline compute engine. You can modify this setting as needed. To enable Real-time R&D, see Tenant Settings. |
|
Apache Flink |
A distributed processing engine for stateful computations over unbounded and bounded data streams. |
|
|
FusionInsight Flink |
A stream processing engine based on Apache Flink for real-time analysis of high-speed data streams. |
|
|
Blink exclusive |
Alibaba Cloud's real-time compute engine. Important
This version is no longer sold on the public cloud. Use with caution. |
|
Single-tenant multi-engine
Tenants created in Dataphin V6.1 or later use single-tenant multi-engine mode by default. You cannot configure the compute engine here. Instead, go to Plan > Cluster Management to create clusters with different engine types. For more information, see Create a cluster in single-tenant multi-engine mode.