All Products
Search
Document Center

Realtime Compute for Apache Flink:Model DDLs

Last Updated:Sep 18, 2026

This document describes the data definition language (DDL) statements for registering, querying, modifying, and deleting AI models in Flink SQL.

Usage notes

Supported only in VVR 11.7 or later. Requires activation of Flink AI Service (built-in models).

Register a model

After the master account activates Flink AI Service in the target region, you can create a model in built-in model mode.

CREATE TEMPORARY MODEL my_llm
INPUT (prompt String COMMENT 'Input prompt')
OUTPUT (response String COMMENT 'Model output')
WITH (
  'provider' = 'dashscope',
  'task' = 'chat/completions',
  'model' = 'qwen3.5-flash'
);
  • The endpoint parameter is not required. The system automatically selects the appropriate service endpoint.

  • The api-key parameter is not required. The system authenticates using the Flink-managed API Key.

  • The task parameter is required to declare the model task type.

WITH parameters

General

Parameter

Description

Data type

Required

Default value

Remarks

provider

The model service type.

String

Yes

None

Valid values: openai-compat, dashscope, and triton.

task

The model task type.

String

Yes

None

Valid values:

model

The specific model to invoke on the service side.

String

Yes

None

Select a model based on the task type. For details, see Built-in model list .

max-context-size

The maximum context size for a single request.

Integer

No

None

If the maximum capacity is exceeded, the action defined in context-overflow-action is triggered.

context-overflow-action

The action to take when a request's context exceeds the maximum capacity.

String

No

truncated-tail

Valid values:

  • truncated-tail: Automatically truncates tokens when the capacity is exceeded, retaining the most recent max-context-size tokens. No logs are recorded.

  • truncated-tail-log: Automatically truncates tokens from the tail that exceed the capacity, retaining the most recent max-context-size tokens. The truncation is logged.

  • truncated-head: Truncates the earliest tokens from the head, retaining the most recent max-context-size tokens.

  • truncated-head-log: Trims the earliest tokens from the head, keeping the latest max-context-size tokens. Logs the truncation.

  • skipped: Discards the data record directly. No logs are recorded.

  • skipped-log: Discards the data and records a log.

error-handling-strategy

The strategy for handling model request errors.

String

No

retry

Valid values:

  • retry: Resend the request.

  • failover: Throw an exception.

  • ignore: Ignore the exception and skip the data record.

retry-num

The number of retry attempts.

int

No

100

Takes effect only when error-handling-strategy = retry.

retry-fallback-strategy

The fallback strategy to use after the maximum number of retries is reached.

String

No

failover

Valid values:

Takes effect only when error-handling-strategy is set to a value other than retry.

retry-backoff-strategy

The retry backoff strategy. Defines how the interval between retries is calculated.

String

No

fixed

Valid values:

  • fixed: Fixed interval.

  • exponential: Exponential interval.

retry-backoff-base-interval

The base time interval for the retry backoff.

Duration

No

1 s

-

content-type

The type of input data. Applies to Chat/Completions and Multimodal-Embedding tasks.

String

No

text

  • Content type for a single input column. Valid values: text (default) or image_url.

  • Mutually exclusive with content-types.

content-types

The content types of the input columns for multimodal models. Applies to Chat/Completions and Multimodal-Embedding tasks.

String

No

None

  • A semicolon-separated list. Each input column corresponds to one type (for example, text;image_url).

    Supported types include text, image_url, and multi_image_urls.

  • Mutually exclusive with content-type.

Note

Supported only in VVR 11.8 or later. For details, see General invocation.

chat/completions

Text generation tasks use chat/completions. The following parameters are supported:

Parameter

Description

Data type

Required

Default value

Remarks

system-prompt

The system prompt for the request.

String

No

"You are a helpful assistant."

An empty string is supported.

temperature

Controls the smoothness of the probability distribution for each candidate token.

Float

No

None

Valid range: [0, 2). A value of 0 is not recommended and has no practical meaning.

A higher temperature flattens the probability distribution, so more low-probability tokens are selected and the output is more diverse. A lower temperature sharpens the distribution, so high-probability tokens are more likely and the output is more deterministic.

top-p

The probability threshold for nucleus sampling.

Float

No

None

A higher value increases randomness, while a lower value increases determinism.

stop

A stop sequence.

String

No

None

The model stops generating content when the specified string is about to be produced.

max-tokens

The maximum number of tokens that the model can generate.

Integer

No

None

Limited by the model's capabilities.

presence-penalty

Controls token repetition.

Double

No

None

Valid range: -2.0 to 2.0. Positive values penalize tokens that have already appeared in the text, making the model more likely to discuss new topics.

n

The number of outputs to generate for each input.

int

No

None

-

seed

A random number seed for the model's response.

Long

No

None

If specified, the model platform attempts deterministic sampling, so repeated requests with the same seed and parameters return the same result whenever possible.

response-format

The format of the return value.

String

No

text

Valid values:

  • text

  • json_object

extra-body

Extra HTTP body for the request.

String

No

None

Must be a JSON-formatted string. For details, see extra-body description.

user-prompt

The user prompt for the request.

String

No

None

Similar to system-prompt, but submitted from the user's role.

embeddings

Text embedding tasks use embeddings. The following parameters are supported:

Parameter

Description

Data type

Required

Default value

Remarks

dimension

The dimension of the output vectors.

Integer

No

None

The supported dimensions depend on the specific model. Common values are 1024, 768, and 512.

multimodal-embedding

Multimodal embedding tasks use multimodal-embedding to convert text, image, or mixed text-and-image inputs into vectors. The following parameters are supported:

Parameter

Description

Data type

Required

Default value

Remarks

dimension

The dimension of the output vectors.

Integer

No

None

The supported dimensions depend on the specific model. Common values are 1024, 768, and 512.

Example 1: Images only

CREATE TEMPORARY MODEL multimodal_embedding_model
INPUT (`input` STRING)
OUTPUT (`content` ARRAY<FLOAT>)
WITH (
  'provider' = 'dashscope',
  'task' = 'multimodal-embedding',
  'model' = 'qwen3-vl-embedding',
  'dimension' = '512',
  'content-type' = 'image_url'
);

Example 2: Text and image

Text-and-image fusion applies only to models whose fusion capability is enabled by enable_fusion, as described in Multimodal fused vectors. Flink determines whether to enable fusion based on the number of input columns of the model. No additional parameter is required.

CREATE TEMPORARY MODEL fusion_embedding_model
INPUT (text_input STRING, image_input STRING)
OUTPUT (embedding ARRAY<FLOAT>)
WITH (
  'provider' = 'dashscope',
  'task' = 'multimodal-embedding',
  'model' = 'qwen3-vl-embedding',
  'dimension' = '512',
  'content-types' = 'text;image_url'
);

extra-body description

The value of extra-body is a JSON-formatted string used to add extra parameters to the model request body. The available parameters depend on the model service provider. The following are common parameters supported by Alibaba Cloud Model Studio (including but not limited to):

Parameter

Type

Description

top_k

Integer

The sampling candidate set size. A larger value increases randomness. For details, see Request body.

enable_thinking

Boolean

Whether to enable deep thinking mode (applies to models that support thinking, such as Qwen3). For details, see Deep thinking.

thinking_budget

Integer

The maximum token length for the thinking process. Takes effect only when enable_thinking is true . For details, see Limit thinking length.

translation_options

Object

Translation parameters for translation models. For details, see Translation capability.

enable_search

Boolean

Whether to use internet search results to assist answer generation. Defaults to false . For details, see Web search.

search_options

Object

Web search strategy configuration. Takes effect only when enable_search is true . For details, see Set search scale strategy.

Example

CREATE MODEL my_model
USING openai_compatible
WITH (
  'provider' = 'openai-compat',
  'model' = 'qwen3.5-flash',
  'task' = 'chat/completions',
  'extra-body' = '{"enable_thinking": true, "thinking_budget": 4096}'
);

Querying models

In the Data Query editor, run one of the following commands.

  • List the names of registered models:

    SHOW MODELS [ ( FROM | IN ) [catalog_name.]database_name ];
  • Show the statement used to create a model:

    SHOW CREATE MODEL [catalog_name.][db_name.]model_name;
  • Show the input and output schema of a model:

    DESCRIBE MODEL [catalog_name.][db_name.]model_name;

Example

SHOW MODELS;

-- RESULT
--+------------+
--| model name |
--+------------+
--|          m |
--+------------+

DESCRIBE MODEL m;

-- RESULT
-- +---------+--------+------+----------+
-- |    name |   type | null | is input |
-- +---------+--------+------+----------+
-- | content | String | TRUE |     TRUE |
-- |   label | BIGINT | TRUE |    FALSE |
-- +---------+--------+------+----------+

Modifying models

In the Data Query editor, run the following command.

ALTER MODEL [IF EXISTS] [catalog_name.][db_name.]model_name {
  RENAME TO new_table_name
  SET (key1=val1, ...)
  RESET (key1, ...)
}

Examples

  • Rename a registered model:

    ALTER MODEL m RENAME TO m1; -- Renames the model to m1.
  • Modify a model parameter:

    ALTER MODEL m SET ('endpoint' = '<Your_Endpoint>'); -- Adjusts the endpoint path.
  • Reset a model parameter to its default value:

    ALTER MODEL m RESET ('endpoint'); -- Resets the endpoint path.

Deleting models

In the Data Query editor, run the following command.

DROP [TEMPORARY] MODEL [IF EXISTS] [catalog_name.][db_name.]model_name

Example

DROP MODEL m;