All Products
Search
Document Center

Dataphin:Activate a semi-managed Dataphin instance

Last Updated:Jun 02, 2026

Activate Dataphin to start using its data development and governance features.

Purchase notes

  • Financial Cloud accounts cannot purchase Dataphin instances.

  • Dataphin supports automatic renewal (configured in the software fee section). Auto-renewal applies only to the Dataphin software — renew underlying resources separately to avoid service disruptions. The cycle is monthly and can be canceled at any time. Renew a semi-managed instance.

  • Service activation takes 2–3 hours after purchase. If activation fails, contact the Dataphin O&M and deployment team.

  • You can purchase and combine add-on modules. Billing.

Important notes

Before activating Dataphin:

  • Contact Alibaba Cloud pre-sales support to confirm Dataphin meets your needs. The support team grants purchase permissions after confirmation.

  • Dataphin does not offer unconditional refunds. Confirm your Dataphin version before purchasing.

    To request a refund due to special circumstances after purchase, submit a ticket and contact your account manager. Refunds are not granted for non-product issues. Approved refunds deduct applicable usage fees.

  • Dataphin uses subscription billing only.

  • Expired instances within the 15-day retention period can only be renewed.

Purchase a Dataphin instance

  1. Log on to Alibaba Cloud with your primary account.

  2. Hover over Products > Big Data Computing > Data Development and Services, then click Intelligent Data Construction and Governance Dataphin.

  3. On the Dataphin product page, click Console or Activate Now (Semi-managed Edition).

  4. On the purchase page, configure the instance name, region, billing method, subscription duration, and optional add-on modules.

    Parameter

    Description

    Service Instance Name

    Enter a name for the Dataphin instance. This name appears on the Dataphin console. You cannot change the name after the instance is created.

    The name can contain digits, letters, hyphens (-), and underscores (_), and must not exceed 64 characters in length.

    The name must follow the service's resource naming conventions.

    Region

    Select a region for the Dataphin instance. Supported regions:

    • China: China (Qingdao), China (Beijing), China (Hohhot), China (Ulanqab), China (Hangzhou), China (Shanghai), China (Shenzhen), China (Heyuan), China (Guangzhou), China (Wuhan-Local Region), China (Chengdu), and China (Hong Kong).

    • Asia Pacific: Japan (Tokyo), South Korea (Seoul), Singapore, Malaysia (Kuala Lumpur), Indonesia (Jakarta), and Thailand (Bangkok).

    • Europe and Americas: US (Virginia), UK (London), and Germany (Frankfurt).

    • Middle East & India: UAE (Dubai).

    Software fee configuration

    Billing method

    Only the subscription billing method is supported.

    Subscription duration

    The system selects 1 Month by default. Supported durations include 1 Month, 1 Year, 2 Years, and 3 Years.

    Auto-renewal

    If you enable this feature, only the Dataphin software is renewed. You must renew the underlying resources separately.

    Dataphin version

    The system automatically fills in the latest version number.

    Dataphin feature selection

    Data processing unit

    The default value is 500. Select a specification. Options include 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, and 10000.

    Real-time development (Optional)

    One-stop, high-performance real-time big data processing with a professional development environment for stream data use cases.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Intelligent O&M (Optional)

    Includes baseline monitoring and throttling configuration to ensure timely data output and system stability while reducing manual O&M costs.

    The system provides 3 Baselines + 1 Throttling Rule (Free) by default. If you need higher specifications, select the Standard Edition.

    Data standard (Optional)

    Create and manage data standards, manage reference data, and associate standards with asset metadata. Combined with quality monitoring, enables end-to-end asset governance across pre-event, in-event, and post-event stages.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Asset quality (Optional)

    Monitors data quality from physical and logical perspectives. Configure quality rules, execute verification tasks, and view comprehensive reports to identify quality risks and rule coverage gaps.

    You can select the Domain Edition or Domain Edition + Global Edition. If you do not need this feature at the moment, select Not Selected.

    Asset security (Optional)

    Define security levels and business categories for data, build identification rules, and set desensitization rules for sensitive data.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Resource governance (Optional)

    Metadata collection, unified metadata model creation, and governance-driven closed-loop processes. Helps asset managers and developers build a manageable data asset health system.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    OpenAPI

    External entry point for sharing data, features, and services across applications. Covers development, O&M, assets, and platform management modules.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Data service (Optional)

    Unifies data thematic units with self-service API configuration, debugging, release, and call monitoring. Enables field-level access control to simplify data consumption while securing open data.

    The available specification is api.base (max 500 QPS/50 concurrency). If you do not need this feature at the moment, select Not Selected.

    Note
    • QPS: Queries Per Second. The average number of requests that can be processed per second.

    • concurrency: The number of API requests that can be handled at the same time.

    QPS, concurrency, and response time (RT) are closely related. With a fixed concurrency, a longer API response time results in a lower QPS. The relationship can be expressed by the following formula:

    QPS = concurrency / RT (in seconds).

    Real-time integration (Optional)

    Supports various scenarios, including incremental integration from data sources like MySQL, Oracle, and PostgreSQL to destinations such as Hive, Kafka, DataHub, and MaxCompute, as well as synchronization from Microsoft SQL Server and IBM DB2 to Kafka.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Gateway type

    The system currently only supports the Dataphin self-developed gateway.

    Tag Factory (Optional)

    Visual tag processing with offline and real-time tag development to improve efficiency and lower the development barrier.

    Supported specifications include Offline Edition, Real-time Edition, Offline Edition + Group Selection, Offline Edition + Real-time Edition, and Offline Edition + Group Selection + Group Permissions. If you do not need this feature at the moment, select Not Selected.

    Row-level permission (Optional)

    Row-level permission control for tables in compute and data sources, governing data visibility for individual accounts, user groups, and production accounts.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Number of Tenants

    Allows you to create multiple tenants within a single Dataphin instance, where different tenants can rely on different computing engines. You can select to create 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 tenants.

    Register Scheduling Cluster (Optional)

    Connects to databases in other networks through a registered scheduling cluster, avoiding cross-network data transfer. Used in hybrid on-premises/cloud scenarios.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Metadata collection (Optional)

    Supports automatic collection of metadata from big data storage engines, such as Hive, StarRocks, and Hologres.

    Supported specifications are Default Version and Big Data Engine.

    Metadata management (Optional)

    Supports enriching and managing object properties and building metamodels for different objects.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Asset operation (Optional)

    Asset publishing and catalog management. Group assets by topic, configure visibility, and simplify asset discovery and consumption.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Asset consumption (Optional)

    Centrally managed data permissions. After applying for permissions in the asset catalog, navigate directly to a BI platform for analysis — no need to create separate data sources or datasets.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Note

    You must purchase the asset operation feature to use asset consumption.

    X-O&M Assistant (Optional)

    Mobile O&M assistant for checking instance status, diagnosing anomalies, and receiving repair suggestions.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Coding Assistant (Optional)

    Intelligent code completion, error correction, and context-based code generation.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Analysis (Optional)

    Turns natural-language questions into SQL queries via Intelligent Analysis. Run queries and get results with one click.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Data Engineering (Optional)

    Automates data integration, subject-domain modeling, conceptual modeling, logical modeling, data processing, and metric generation.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Data Standard (Optional)

    Uses LLM capabilities to extract code tables and standard definitions, and intelligently recommends field-to-standard mappings.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Data Security (Optional)

    Uses LLM to identify sensitive data and recommend classification mappings.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Asset Q&A (Optional)

    Natural-language search for data assets based on a knowledge base.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Directory Management (Optional)

    Uses LLM to enrich attributes of table and metric assets.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-Data Quality (Optional)

    Traces quality issue root causes via sampled data and lineage analysis. Generates remediation suggestions and impact assessments to close the governance loop.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    X-App Creation (Optional)

    Quickly build micro apps from Dataphin data service APIs to accelerate data value delivery.

    You can select the Standard Edition. If you do not need this feature at the moment, select Not Selected.

    Dataphin domain name settings

    Product access domain name

    The domain name used to access the Dataphin product instance. It cannot be the same as the OpenAPI Access Domain Name or the Data Service Access Domain Name. Example: dataphin.yourcompany.com.

    It can contain lowercase letters, digits, hyphens (-), and periods (.). Hyphens cannot be used alone, consecutively, or at the beginning or end. The domain name must be 63 characters or fewer.

    OpenAPI access domain name

    This parameter is available only when the OpenAPI feature is enabled.

    The domain name for accessing the Dataphin product instance through OpenAPI. It cannot be the same as the Product Access Domain Name or the Data Service Access Domain Name. Example: dataphin-openapi.yourcompany.com.

    The naming rules are the same as for the Product Access Domain Name.

    Data service access domain name

    This parameter is available only when the data service feature is enabled.

    The domain name that points to the data service application. It cannot be the same as the Product Access Domain Name or the OpenAPI Access Domain Name. Example: dataphin-dataservice.yourcompany.com.

    The naming rules are the same as for the Product Access Domain Name.

    Enable public access

    • If enabled, the system automatically creates a pay-by-traffic EIP for the Ingress LB instance. After you bind the EIP and domain name in your local hosts file, you can access the Dataphin instance from outside your office network.

    • If you want to restrict access to the Dataphin instance to only your office network, disable this feature.

    Key pair configuration

    Key pair name

    The key pair used to log on to ECS instances.

    Network configuration

    Zone 1, Zone 2

    Dataphin deploys its underlying resources across multiple zones for disaster recovery.

    VPC ID

    Select the VPC in which you want to deploy the Dataphin instance. Choose carefully because this setting cannot be changed after creation.

    vSwitch ID 1, vSwitch ID 2

    Select two vSwitches in the selected VPC.

    Kubernetes version

    Use the system default configuration.

    Pod Network CIDR, Service CIDR

    • Dataphin depends on Container Service for Kubernetes (ACK) as an underlying resource. The network plug-in used during ACK deployment is Flannel, which requires you to specify CIDR blocks. Comparison of Terway and Flannel.

    • The pod and Service CIDR blocks are virtual network segments. They must not overlap with each other or with the VPC's vSwitch CIDR blocks. For example, if the VPC network segment is 172.16.0.0/12, the Kubernetes pod address range cannot be 172.16.0.0/16 or 172.17.0.0/16, because these ranges are included in 172.16.0.0/12.

    • The number of reserved IP addresses for the pod CIDR and Service CIDR affects the concurrency of Dataphin tasks. We recommend reserving 2,048 or more IP addresses (a CIDR prefix of /21 or shorter). Choose carefully, because this setting cannot be changed after creation. Flannel network mode.

    Automatically configure NAT gateway

    Allows cluster nodes and applications to access the public internet.

    • If enabled and the selected VPC already has a NAT gateway, ACK uses that gateway by default and automatically configures SNAT rules. If the selected VPC does not have a NAT gateway, ACK automatically creates one and configures SNAT rules.

    • If you disable this option, you must ensure the ACK cluster has public internet access. Dataphin requires this access to pull images during deployment; otherwise, the deployment will fail.

    Advanced configurations

    Application node pool instance type

    Select an appropriate node type and number of nodes based on the application deployment mode. Avoid mixing different specifications. Supported types include 16 vCPUs/128 GiB (ecs.r9i.4xlarge), 16 vCPUs/128 GiB (ecs.r8i.4xlarge), 16 vCPUs/128 GiB (ecs.u2i-c1m8.4xlarge), 16 vCPUs/128 GiB (ecs.r7.4xlarge), 16 vCPUs/128 GiB (ecs.u1-c1m8.4xlarge), and 16 vCPUs/128 GiB (ecs.hfr7.4xlarge).

    Note

    The maximum resource requirement for high-availability mode is 40 vCPUs/320 GiB, while for non-high-availability mode it is 20 vCPUs/160 GiB.

    Initial number of nodes in application node pool

    Configure the initial number of nodes based on the application deployment mode. We recommend 3, with a minimum of 2. If you set the number to 2, high availability is not guaranteed.

    Scheduling node pool instance type

    Configure the node pool instance type based on the number of scheduling tasks. Supported types include 24 vCPUs/96 GiB (ecs.g9i.6xlarge), 24 vCPUs/96 GiB (ecs.g8i.6xlarge), 24 vCPUs/96 GiB (ecs.g7.6xlarge), 24 vCPUs/96 GiB (ecs.hfg7.6xlarge), and 24 vCPUs/96 GiB (ecs.hfg6.6xlarge).

    Initial number of nodes in scheduling node pool

    Configure the initial number of nodes based on the number of scheduling tasks. We recommend 2, with a minimum of 1. If you set the number to 1, high availability is not guaranteed.

    Number of Dataphin application deployment replicas

    Ensure that the application node pool has sufficient resources before you increase the number of replicas. The default is 2.

    PostgreSQL specifications

    Select the maximum number of connections for the PostgreSQL database. Supported specifications are 4 vCPUs/16 GiB (max 1600 connections), 8 vCPUs/16 GiB (max 1600 connections), and 16 vCPUs/32 GiB (max 3200 connections).

    Initial disk size of PostgreSQL database (GB)

    Primarily used to store scheduling instances and business metadata. The default is 400 GB. You can configure a storage space between 200 GB and 400 GB, in 5 GB increments.

  5. Carefully review the purchase information. If it is correct, click Next: Confirm Order.

  6. On the Confirm Order page, verify the instance specifications. Click the Intelligent Data Construction and Governance Service Agreement link next to Service Agreement and read the terms. Select I have read and agree to the Intelligent Data Construction and Governance Service Agreement, then click Pay.

    image.png

Next steps

After activation, obtain the initial Alibaba Cloud account and IP address, then bind the hostname to the IP in your hosts file. Perform a cold start after deployment.