Create an Amazon S3 data source to enable Dataphin to read business data from or write data to Amazon S3.
Background information
Amazon S3 (Simple Storage Service) is a cloud storage service provided by Amazon for storing and retrieving data in the cloud. To develop data in Dataphin or write Dataphin data to Amazon S3, you must first create an Amazon S3 data source. For more information, see What is Amazon S3.
Permission requirements
Only custom global roles with the Create Data Source permission and the system roles super administrator, data source administrator, domain architect, and project administrator can create data sources.
Procedure
-
On the Dataphin homepage, click Management Hub > Datasource Management in the top navigation bar.
-
On the Datasource page, click +Create Data Source.
-
In the Create Data Source page, select Amazon S3 from the File section.
You can also select Amazon S3 from the Recently Used section or search for it by entering keywords in the search box.
-
On the Create Amazon S3 Data Source page, configure the connection parameters.
-
Configure the basic information of the data source.
Parameter
Description
Datasource Name
Enter the name of the data source. The name must meet the following requirements:
-
It can contain only Chinese characters, letters, digits, underscores (_), and hyphens (-).
-
It cannot exceed 64 characters in length.
Datasource Code
After you configure the data source code, you can reference tables in the data source in Flink_SQL tasks by using the format
data_source_code.table_nameordata_source_code.schema.table_name. If you need to automatically access the data source in the corresponding environment based on the current environment, use the variable format${data_source_code}.tableor${data_source_code}.schema.table. For more information, see Dataphin data source table development method.Important-
The data source code cannot be modified after it is configured.
-
You can preview data on the object details page in the asset directory and asset checklist only after the data source code is configured.
-
In Flink SQL, only MySQL, Hologres, MaxCompute, Oracle, StarRocks, Hive, SelectDB, and GaussDB data warehouse service (DWS) data sources are currently supported.
Data Source Description
A brief description of the data source. It cannot exceed 128 characters.
Data Source Configuration
Select the data source to configure:
-
If the business data source distinguishes between production and development data sources, select Production + Development Data Source.
-
If the business data source does not distinguish between production and development data sources, select Production Data Source.
Tag
Categorize data sources using tags. For information on how to create tags, see Manage data source tags.
-
-
Configure the connection parameters between the data source and Dataphin.
If you select Production + Development data source, configure connection information for both environments. If you select Production data source, configure connection information for the production environment only.
NoteFor environment isolation, configure production and development as separate data sources. However, Dataphin also supports using the same data source with identical parameter values for both.
Parameter
Description
Endpoint
The endpoint of the region where Amazon S3 is located. The format is
http://s3-{Region}.amazonaws.com, where Region is the region where the bucket is located.Each region uses a different endpoint. For more information, see Amazon S3 endpoints.
Region
The region where the bucket is located. Required only if the region is not specified in the endpoint.
Bucket
The Amazon S3 bucket for storing objects. For more information, see Amazon S3 bucket overview.
Directory
If you only have permissions for a specific directory, you can specify the directory path here. For example,
/dataphin/.Access ID, Access Key
The AccessKey ID and AccessKey Secret of the account where the Amazon S3 data source is located.
To obtain these credentials, see Amazon access keys.
NoteThese are not Alibaba Cloud account AccessKey ID and AccessKey Secret.
-
-
Select a Default Resource Group, which will be used to run tasks related to the current data source, including database SQL, offline database migration, and data preview.
-
Click Test Connection or directly click OK to save and complete the creation of the Amazon S3 data source.
Clicking Test Connection tests whether Dataphin can connect to the data source. If you click OK directly, the system automatically tests the connection for all selected clusters. The data source is created even if the connection test fails.