This topic describes the major feature updates and bug fixes for Realtime Compute for Apache Flink that were released on March 4, 2022.
Overview
On March 4, 2022, we released Ververica Runtime (VVR) 4.0.12, which is based on Apache Flink 1.13. This new version supports adaptive JSON schema evolution for common Kafka->Flink->Hologres pipelines. For Data Lake Formation, we released an enterprise-grade Hudi connector. To improve development efficiency, we provide more than 20 common Flink SQL job templates. To enhance operations and maintenance (O&M) services, we provide powerful job diagnostics and the ability to dynamically adjust log levels without stopping jobs. The release also includes many powerful data processing features, such as enterprise-grade features for ClickHouse, new connectors, and new syntax for data warehousing and data lake ingestion. In addition, this new version fixes several bugs that were resolved in the open-source Apache Flink community.
New features
Feature | Details | References |
Adaptive JSON schema evolution for Hologres | JSON is one of the most common event formats in stream processing. For real-time streaming jobs and tables in the backend storage engine, schema evolution should be a transparent process. This new version includes the following enhancements:
| |
Enhanced capabilities for building Iceberg and Hudi data lakes |
| |
Improved usability for log viewing and settings |
| |
Multiple enterprise-grade features for Flink and ClickHouse |
| |
Optimized job diagnosis rules and interface |
| |
Data synchronization supports adding computed columns | The CTAS statement supports adding a computed column to a source table and setting the new column as the primary key of the sink table. When ingesting data into a data warehouse or data lake, the CTAS statement lets you specify the position of a new computed column and make it a physical column in the sink table. The results of the computed column are synchronized to the sink table in real time. The CTAS statement also supports changing the primary key of the sink table to the new computed column. | |
Easier generation of test data | Added a simulated data generator connector. The simulated data generator connector helps you easily generate test data that is relevant to your business scenarios. This meets your needs for verifying business logic during development and testing. | |
New Template Hub to accelerate job development |
| |
Clearer display of resource usage | In the lower-left corner of the Flink development console, the CPU and memory usage for the current project is displayed. This helps you quickly manage project resources. | None |
Quickly locate logs for slow checkpoint nodes | In the snapshot history, you can now sort node snapshot statuses. You can also go directly to the TM logs from the snapshot history interface with one click to view the cause of slow checkpoints. | Locate slow checkpoints and view the logs of corresponding TaskManagers |
Support for AnalyticDB for PostgreSQL sink tables and dimension tables |
| |
Improved usability of the enterprise-grade state backend |
|
Performance optimizations
The enterprise-grade state backend in this new version includes many optimizations. It greatly improves the performance of two-stream or multi-stream join jobs. The average utilization of compute resources can be increased by 50%, and by 100% to 200% in typical scenarios. This helps you run stateful stream computing applications more smoothly.
Bug fixes
Optimized the Catalog service to resolve an issue where refreshing failed when the number of databases or tables was large.
Fixed an issue where the Flink version was not displayed for session clusters.
Fixed an issue with the display of the WatermarkLag curve on the Metrics page.
Optimized the paging display for curves on the Metrics page.
Fixed issues with Flink CDC, including the currentFetchEventTimeLag metric and class conflicts.
Fixed an issue where the CTAS syntax could not be used to modify existing columns.