This document describes the data definition language (DDL) statements for registering, querying, modifying, and deleting AI models in Flink SQL.
Usage notes
Supported only in VVR 11.7 or later. Requires activation of Flink AI Service (built-in models).
Register a model
After the master account activates Flink AI Service in the target region, you can create a model in built-in model mode.
CREATE TEMPORARY MODEL my_llm
INPUT (prompt String COMMENT 'Input prompt')
OUTPUT (response String COMMENT 'Model output')
WITH (
'provider' = 'dashscope',
'task' = 'chat/completions',
'model' = 'qwen3.5-flash'
);
-
The
endpointparameter is not required. The system automatically selects the appropriate service endpoint. -
The
api-keyparameter is not required. The system authenticates using the Flink-managed API Key. -
The
taskparameter is required to declare the model task type.
WITH parameters
General
|
Parameter |
Description |
Data type |
Required |
Default value |
Remarks |
|
provider |
The model service type. |
String |
Yes |
None |
Valid values:
|
|
task |
The model task type. |
String |
Yes |
None |
Valid values: |
|
model |
The specific model to invoke on the service side. |
String |
Yes |
None |
Select a model based on the task type. For details, see Built-in model list . |
|
max-context-size |
The maximum context size for a single request. |
Integer |
No |
None |
If the maximum capacity is exceeded, the action defined in context-overflow-action is triggered. |
|
context-overflow-action |
The action to take when a request's context exceeds the maximum capacity. |
String |
No |
|
Valid values:
|
|
error-handling-strategy |
The strategy for handling model request errors. |
String |
No |
retry |
Valid values:
|
|
retry-num |
The number of retry attempts. |
int |
No |
100 |
Takes effect only when |
|
retry-fallback-strategy |
The fallback strategy to use after the maximum number of retries is reached. |
String |
No |
failover |
Valid values: Takes effect only when |
|
retry-backoff-strategy |
The retry backoff strategy. Defines how the interval between retries is calculated. |
String |
No |
fixed |
Valid values:
|
|
retry-backoff-base-interval |
The base time interval for the retry backoff. |
Duration |
No |
1 s |
- |
|
content-type |
The type of input data. Applies to Chat/Completions and Multimodal-Embedding tasks. |
String |
No |
text |
|
|
content-types |
The content types of the input columns for multimodal models. Applies to Chat/Completions and Multimodal-Embedding tasks. |
String |
No |
None |
Note
Supported only in VVR 11.8 or later. For details, see General invocation. |
chat/completions
Text generation tasks use chat/completions. The following parameters are supported:
|
Parameter |
Description |
Data type |
Required |
Default value |
Remarks |
|
system-prompt |
The system prompt for the request. |
String |
No |
"You are a helpful assistant." |
An empty string is supported. |
|
temperature |
Controls the smoothness of the probability distribution for each candidate token. |
Float |
No |
None |
Valid range: [0, 2). A value of 0 is not recommended and has no practical meaning. A higher temperature flattens the probability distribution, so more low-probability tokens are selected and the output is more diverse. A lower temperature sharpens the distribution, so high-probability tokens are more likely and the output is more deterministic. |
|
top-p |
The probability threshold for nucleus sampling. |
Float |
No |
None |
A higher value increases randomness, while a lower value increases determinism. |
|
stop |
A stop sequence. |
String |
No |
None |
The model stops generating content when the specified string is about to be produced. |
|
max-tokens |
The maximum number of tokens that the model can generate. |
Integer |
No |
None |
Limited by the model's capabilities. |
|
presence-penalty |
Controls token repetition. |
Double |
No |
None |
Valid range: -2.0 to 2.0. Positive values penalize tokens that have already appeared in the text, making the model more likely to discuss new topics. |
|
n |
The number of outputs to generate for each input. |
int |
No |
None |
- |
|
seed |
A random number seed for the model's response. |
Long |
No |
None |
If specified, the model platform attempts deterministic sampling, so repeated requests with the same seed and parameters return the same result whenever possible. |
|
response-format |
The format of the return value. |
String |
No |
text |
Valid values:
|
|
extra-body |
Extra HTTP body for the request. |
String |
No |
None |
Must be a JSON-formatted string. For details, see extra-body description. |
|
user-prompt |
The user prompt for the request. |
String |
No |
None |
Similar to system-prompt, but submitted from the user's role. |
embeddings
Text embedding tasks use embeddings. The following parameters are supported:
|
Parameter |
Description |
Data type |
Required |
Default value |
Remarks |
|
dimension |
The dimension of the output vectors. |
Integer |
No |
None |
The supported dimensions depend on the specific model. Common values are 1024, 768, and 512. |
multimodal-embedding
Multimodal embedding tasks use multimodal-embedding to convert text, image, or mixed text-and-image inputs into vectors. The following parameters are supported:
|
Parameter |
Description |
Data type |
Required |
Default value |
Remarks |
|
dimension |
The dimension of the output vectors. |
Integer |
No |
None |
The supported dimensions depend on the specific model. Common values are 1024, 768, and 512. |
Example 1: Images only
CREATE TEMPORARY MODEL multimodal_embedding_model
INPUT (`input` STRING)
OUTPUT (`content` ARRAY<FLOAT>)
WITH (
'provider' = 'dashscope',
'task' = 'multimodal-embedding',
'model' = 'qwen3-vl-embedding',
'dimension' = '512',
'content-type' = 'image_url'
);
Example 2: Text and image
Text-and-image fusion applies only to models whose fusion capability is enabled by enable_fusion, as described in Multimodal fused vectors. Flink determines whether to enable fusion based on the number of input columns of the model. No additional parameter is required.
CREATE TEMPORARY MODEL fusion_embedding_model
INPUT (text_input STRING, image_input STRING)
OUTPUT (embedding ARRAY<FLOAT>)
WITH (
'provider' = 'dashscope',
'task' = 'multimodal-embedding',
'model' = 'qwen3-vl-embedding',
'dimension' = '512',
'content-types' = 'text;image_url'
);
extra-body description
The value of extra-body is a JSON-formatted string used to add extra parameters to the model request body. The available parameters depend on the model service provider. The following are common parameters supported by Alibaba Cloud Model Studio (including but not limited to):
|
Parameter |
Type |
Description |
|
|
Integer |
The sampling candidate set size. A larger value increases randomness. For details, see Request body. |
|
|
Boolean |
Whether to enable deep thinking mode (applies to models that support thinking, such as Qwen3). For details, see Deep thinking. |
|
|
Integer |
The maximum token length for the thinking process. Takes effect only when |
|
|
Object |
Translation parameters for translation models. For details, see Translation capability. |
|
|
Boolean |
Whether to use internet search results to assist answer generation. Defaults to |
|
|
Object |
Web search strategy configuration. Takes effect only when |
Example
CREATE MODEL my_model
USING openai_compatible
WITH (
'provider' = 'openai-compat',
'model' = 'qwen3.5-flash',
'task' = 'chat/completions',
'extra-body' = '{"enable_thinking": true, "thinking_budget": 4096}'
);
Querying models
In the Data Query editor, run one of the following commands.
-
List the names of registered models:
SHOW MODELS [ ( FROM | IN ) [catalog_name.]database_name ]; -
Show the statement used to create a model:
SHOW CREATE MODEL [catalog_name.][db_name.]model_name; -
Show the input and output schema of a model:
DESCRIBE MODEL [catalog_name.][db_name.]model_name;
Example
SHOW MODELS;
-- RESULT
--+------------+
--| model name |
--+------------+
--| m |
--+------------+
DESCRIBE MODEL m;
-- RESULT
-- +---------+--------+------+----------+
-- | name | type | null | is input |
-- +---------+--------+------+----------+
-- | content | String | TRUE | TRUE |
-- | label | BIGINT | TRUE | FALSE |
-- +---------+--------+------+----------+
Modifying models
In the Data Query editor, run the following command.
ALTER MODEL [IF EXISTS] [catalog_name.][db_name.]model_name {
RENAME TO new_table_name
SET (key1=val1, ...)
RESET (key1, ...)
}
Examples
-
Rename a registered model:
ALTER MODEL m RENAME TO m1; -- Renames the model to m1. -
Modify a model parameter:
ALTER MODEL m SET ('endpoint' = '<Your_Endpoint>'); -- Adjusts the endpoint path. -
Reset a model parameter to its default value:
ALTER MODEL m RESET ('endpoint'); -- Resets the endpoint path.
Deleting models
In the Data Query editor, run the following command.
DROP [TEMPORARY] MODEL [IF EXISTS] [catalog_name.][db_name.]model_name
Example
DROP MODEL m;