All Products
Search
Document Center

Realtime Compute for Apache Flink:September 19, 2022 release

Last Updated:Aug 19, 2026

This topic describes the major features and bug fixes in the September 19, 2022 release of Realtime Compute for Apache Flink.

Overview

On September 19, 2022, Realtime Compute for Apache Flink released a new version that includes updates to the platform and engine, connector updates, performance optimizations, and bug fixes. This release includes VVR-4.0.15 based on Apache Flink 1.13 and VVR-6.0.2 based on Apache Flink 1.15. The key updates are as follows:

  • This release introduces VVR 6.0.2, the first enterprise-grade Flink engine based on Apache Flink 1.15. This version integrates major features and performance optimizations from the open-source community, including enhancements to window table-valued functions, CAST functions, the type system, and JSON functions.

  • State management has been a major focus for our users. This release unifies checkpoint and savepoint management into a single feature: status set management. This significantly improves the speed of savepoint creation and recovery, reduces savepoint size, and enhances overall success rate and stability.

    Also, savepoints are no longer deleted when a deployment is canceled, which changes the previous behavior. Now, checkpoints and savepoints are distinct, and you can explicitly create and manage savepoints. In addition to improved usability, optimizations in the Gemini state backend lead to significant cost savings. The new status set management can reduce your annual OSS storage costs by 15% to 40%. The platform also allows you to start a deployment from a savepoint created by another deployment, which simplifies A/B testing and other dual-run scenarios.

  • To improve resource utilization, we have introduced scheduled tuning. If your workload has predictable traffic peaks and valleys, you can create a policy to automatically adjust deployment resources to a preset size at a specific time. This helps you manage resource scaling without manual intervention and reduces labor costs.

  • To help diagnose deployments, we are introducing the health score. This feature analyzes deployments throughout their lifecycle, from startup to running, and provides diagnostic information and recommendations to help you maintain your streaming deployments.

  • For better integration, the platform now provides a new set of OpenAPI operations, allowing you to integrate its capabilities into your own services.

  • Real-time risk control is a primary use case for Flink. We previously offered a preview of new features for complex event processing (CEP) on continuous event sequences to select customers, and it has been successfully validated in production.

    In this release, we are making a series of CEP enhancements generally available. First, the highly requested hot-update capability for CEP rules is now available. This allows you to update rules during peak business hours without restarting your deployment, which eliminates the ten-minute service interruptions that risk control systems previously faced during rule updates and significantly improves business availability. Second, we have enhanced the CEP SQL syntax with new extensions to improve its expressiveness. This allows you to convert complex DataStream API deployments into simpler SQL deployments, improving development efficiency and making it easier to integrate with data lineage systems. Finally, this release introduces several new CEP metrics to provide detailed insights into rule matching.

  • Other optimizations include performance improvements. We now automatically enable key-value separation for dual-stream join operators in Flink SQL. This optimization significantly improves the performance of dual-stream join deployments without requiring user configuration. We have also expanded the range of supported Hive versions for Hive Catalog to include 2.1.0-2.3.9 and 3.1.0-3.1.3. For connectors, we have added support for reading from Tablestore and for using the JDBC connector with source tables, dimension tables, and result tables.

New features

Feature

Description

Documentation

Status set management

Status set management decouples state management from deployment start and stop operations for all stateful Flink deployments. Savepoints are no longer deleted when a deployment is stopped. You can use a dedicated management page to create and delete savepoints on a schedule.

Scheduled tuning

For Flink deployments with predictable traffic peaks and valleys, you can define custom scheduling policies. At the specified times, the deployment's resources are automatically adjusted to a preset size to handle traffic fluctuations, eliminating the need for manual scaling.

Health score

The health score feature applies expert rules to detect issues during deployment startup and execution, providing actionable recommendations. This feature helps you better understand the status of your deployments and adjust parameters accordingly.

Perform intelligent deployment diagnostics

Improved member authorization

The authorization process is improved: instead of manually entering user information, you can now select from a list of all RAM users when granting permissions.

Grant permissions on namespaces

Dynamic complex event processing (CEP)

CEP provides pattern matching capabilities for real-time data streams. This release builds on open source Flink CEP by allowing you to externalize deployment rules in a database so they can be loaded dynamically. This is exposed through the DataStream API.

Enhancement of CEP SQL

The MATCH_RECOGNIZE statement allows you to describe CEP rules using SQL. This release enhances the open source Flink MATCH_RECOGNIZE statement with new capabilities, such as outputting timed-out matches and supporting notFollowedBy.

In addition, new metrics have been introduced:

  • patternMatchedTimes: The number of times a pattern was successfully matched.

  • patternMatchingAvgTime: The average time taken for pattern matching.

CEP statements

Support for database synchronization to Kafka

When you use this feature, data is synchronized to a corresponding Upsert Kafka table. You can use the table in Kafka directly instead of the MySQL table, which reduces the load on the MySQL service from multiple deployments.

Define partitioned tables in Hologres result tables with DDL

You can use PARTITION BY to define a partitioned table when you create a Hologres result table.

CREATE TABLE AS statement

Set timeout for asynchronous requests in Hologres dimension tables

By setting the asyncTimeoutMs parameter for asynchronous requests, you can ensure that the application completes data requests within a specific time frame.

Hologres dimension table

Set table properties when creating tables with Hologres Catalog

Setting appropriate table properties can help the system organize and query data efficiently. When you use Hologres Catalog to create a table, you can now set physical table properties in the WITH clause.

Manage Hologres catalogs

MaxCompute sink connector supports the Binary type

  • The Binary data type is now supported. MaxCompute limits the length of this type to 8 MB.

  • The MaxCompute Stream Tunnel Sink feature is added.

  • The flush efficiency of the MaxCompute sink is optimized.

MaxCompute result table

Hive Catalog supports more Hive versions

This version supports Hive 2.1.0-2.3.9 and 3.1.0-3.1.3.

Manage Hive catalogs

Tablestore source connector released

Supports reading incremental logs from Tablestore.

Tablestore source table

JDBC connector released

The community JDBC connector is now built-in.

Parallelism of a Message Queue for Apache RocketMQ source table can exceed the topic partition count

This mode allows you to pre-allocate resources for potential increases in topic partitions before consumption begins.

Message Queue for Apache RocketMQ source table

Set Message Key for Message Queue for Apache RocketMQ result tables

You can now set the message key when writing to Message Queue for Apache RocketMQ.

Message Queue for Apache RocketMQ result table

Support for AnalyticDB for MySQL Catalog

With this catalog, you can directly read metadata from AnalyticDB for MySQL without manually registering AnalyticDB for MySQL tables. This improves development efficiency and ensures data correctness.

Manage AnalyticDB for MySQL catalogs

Performance optimizations

  • This release introduces a native savepoint format, which resolves timeout issues that previously occurred with standard format savepoints for deployments with large states. This significantly improves overall deployment stability.

    Metric

    Improvement

    Savepoint completion time

    An average improvement of 5 to 10 times, with the ratio increasing as the incremental state size decreases. In some typical deployments, the improvement can be up to 100 times.

    Deployment recovery time

    An average improvement of about 5 times, with the ratio increasing as the state size grows.

    Savepoint space overhead

    An average space overhead reduction of 2 times, with the ratio increasing as the state size grows.

    Savepoint network overhead

    An average network overhead reduction of 5 to 10 times, with the ratio increasing as the incremental state size decreases.

  • The dual-stream join operator now automatically infers when to enable key-value separation to optimize performance. For SQL deployments, the dual-stream join operator automatically analyzes deployment characteristics and enables key-value separation to optimize performance. In performance tests for typical scenarios, the average performance improved by more than 40%. For more information, see Optimize high-performance Flink SQL and Configure enterprise-grade state backends.

  • Deployment startup is now 15% faster on average.

Bug fixes

  • Fixed an issue where the modification time of a deployment was updated incorrectly.

  • Fixed an issue where the state of some deployments could not be determined after being suspended and restarted.

  • Fixed an issue where JAR files could not be uploaded locally from Alibaba Finance Cloud.

  • Fixed an issue where the total resources used by a running deployment were inconsistent with the statistics on the page.

  • Fixed an issue where page navigation failed in deployment diagnostic logs.

  • Fixed an error when reading an upsert Kafka table directly from a Kafka Catalog.

  • Fixed a NullPointerException when using intermediate results in nested operations with multiple user-defined functions (UDFs).

  • Fixed issues in mysql-cdc, including abnormal chunk splitting, out-of-memory (OOM) errors, and inconsistent time zones between initial and incremental data. For more information, see MySQL CDC source table.