All Products
Search
Document Center

Realtime Compute for Apache Flink:March 4, 2022 release

Last Updated:Aug 19, 2026

This topic describes the major feature updates and bug fixes for Realtime Compute for Apache Flink that were released on March 4, 2022.

Overview

On March 4, 2022, we released Ververica Runtime (VVR) 4.0.12, which is based on Apache Flink 1.13. This new version supports adaptive JSON schema evolution for common Kafka->Flink->Hologres pipelines. For Data Lake Formation, we released an enterprise-grade Hudi connector. To improve development efficiency, we provide more than 20 common Flink SQL job templates. To enhance operations and maintenance (O&M) services, we provide powerful job diagnostics and the ability to dynamically adjust log levels without stopping jobs. The release also includes many powerful data processing features, such as enterprise-grade features for ClickHouse, new connectors, and new syntax for data warehousing and data lake ingestion. In addition, this new version fixes several bugs that were resolved in the open-source Apache Flink community.

New features

Feature

Details

References

Adaptive JSON schema evolution for Hologres

JSON is one of the most common event formats in stream processing. For real-time streaming jobs and tables in the backend storage engine, schema evolution should be a transparent process.

This new version includes the following enhancements:

  • The table schema can be defined based on the JSON schema before the data is consumed.

  • If the JSON schema changes during data consumption, the schema of the backend Hologres table also changes accordingly.

Enhanced capabilities for building Iceberg and Hudi data lakes

  • Supports Alibaba Cloud Data Lake Formation (DLF) as a catalog.

    A DLF catalog lets you access Hudi, Iceberg, and other DLF-supported engines to quickly build a real-time data lake.

  • A built-in, enterprise-grade Hudi connector for Realtime Compute for Apache Flink reduces O&M complexity.

    • Supports ingesting entire databases into a data lake using Flink CDC and automatically synchronizing table schema changes.

    • Integrates with components such as Alibaba Cloud OSS and DLF to improve data connectivity between compute engines.

Improved usability for log viewing and settings

  • Added log paging.

    On the Job Explorer tab, log paging prevents the page from failing to open when logs become too large for long-running jobs.

  • Support for dynamic log level modification.

    Without restarting a job, you can dynamically modify the log level of a running TaskManager (TM) on the Job Explorer tab to help with troubleshooting.

  • Support for viewing logs of failed TMs.

    On the Job Explorer tab, you can view the logs of TMs that have failed while the JobManager (JM) is still running. This helps you investigate the cause of TM failures.

Multiple enterprise-grade features for Flink and ClickHouse

  • Support for exactly-once semantics.

    Provides exactly-once semantics for the ClickHouse in the open-source big data platform E-MapReduce component, not the cloud-based ClickHouse product.

  • Support for the ClickHouse Nested type.

    The ClickHouse Nested type can be mapped to the Flink Array type.

  • Support for writing directly to local tables of a ClickHouse distributed table.

    Writing directly to the local tables of a distributed table can significantly increase the write throughput for the ClickHouse distributed table.

ClickHouse sink tables

Optimized job diagnosis rules and interface

  • Added more than 20 new diagnosis rules to comprehensively analyze the running status of jobs.

    Provides high, medium, and low risk level alerts based on the actual status of the job.

  • The diagnosis interface is optimized to help you better identify issues.

Intelligent job diagnosis

Data synchronization supports adding computed columns

The CTAS statement supports adding a computed column to a source table and setting the new column as the primary key of the sink table.

When ingesting data into a data warehouse or data lake, the CTAS statement lets you specify the position of a new computed column and make it a physical column in the sink table. The results of the computed column are synchronized to the sink table in real time. The CTAS statement also supports changing the primary key of the sink table to the new computed column.

CREATE TABLE AS (CTAS) statements

Easier generation of test data

Added a simulated data generator connector.

The simulated data generator connector helps you easily generate test data that is relevant to your business scenarios. This meets your needs for verifying business logic during development and testing.

New Template Hub to accelerate job development

  • Provides more than 20 code templates.

    More than 20 templates for common Flink SQL scenarios help you quickly learn how to write job code using Flink SQL.

  • Provides a data synchronization template for MySQL to Hologres.

    Helps you quickly create Flink CDC jobs to complete data synchronization for warehousing and data lake ingestion.

Clearer display of resource usage

In the lower-left corner of the Flink development console, the CPU and memory usage for the current project is displayed. This helps you quickly manage project resources.

None

Quickly locate logs for slow checkpoint nodes

In the snapshot history, you can now sort node snapshot statuses. You can also go directly to the TM logs from the snapshot history interface with one click to view the cause of slow checkpoints.

Locate slow checkpoints and view the logs of corresponding TaskManagers

Support for AnalyticDB for PostgreSQL sink tables and dimension tables

  • Flink supports writing data to AnalyticDB for PostgreSQL sink tables.

  • Flink supports joining with AnalyticDB for PostgreSQL for lookup queries.

Improved usability of the enterprise-grade state backend

  • Added the ability to optimize parameters in real time. This minimizes the complexity and cost of manual tuning and can eliminate the need for manual parameter tuning in more than 95% of cases.

  • Single-core throughput is increased by 10% to 40%. This helps you easily handle fluctuating scenarios such as traffic peaks and troughs.

Performance optimizations

The enterprise-grade state backend in this new version includes many optimizations. It greatly improves the performance of two-stream or multi-stream join jobs. The average utilization of compute resources can be increased by 50%, and by 100% to 200% in typical scenarios. This helps you run stateful stream computing applications more smoothly.

Bug fixes

  • Optimized the Catalog service to resolve an issue where refreshing failed when the number of databases or tables was large.

  • Fixed an issue where the Flink version was not displayed for session clusters.

  • Fixed an issue with the display of the WatermarkLag curve on the Metrics page.

  • Optimized the paging display for curves on the Metrics page.

  • Fixed issues with Flink CDC, including the currentFetchEventTimeLag metric and class conflicts.

  • Fixed an issue where the CTAS syntax could not be used to modify existing columns.