All Products
Search
Document Center

Realtime Compute for Apache Flink:Alibaba Cloud Fluss for Apache Fluss

Last Updated:Aug 27, 2026

Alibaba Cloud Fluss for Apache Fluss is a high-performance, columnar stream storage system built on Apache Fluss. It delivers millisecond-level read and write responses, real-time data updates, and partial column updates. Fluss also provides a Change Data Capture (CDC) log subscription feature, which ensures consistent reads of both full and incremental data for real-time data consumption scenarios.

Prerequisites

Before you begin, it helps to understand the following concepts:

What is Apache Fluss?

What is Apache Paimon?

Product overview

The lake-stream integration architecture of Alibaba Cloud Fluss for Apache Fluss enables seamless integration with data lake formats like Apache Paimon. This approach not only ensures high data timeliness but also significantly reduces the cost of building and maintaining a real-time data warehouse. Through its unified data flow and lake-stream collaboration, Fluss helps organizations break down data silos and accelerate value discovery and efficient data sharing within the data lake.

Features

Description

Millisecond-level stream reads and writes

Fluss supports low-latency streaming read and write operations. You can use it with Apache Flink to build a high-throughput, low-latency real-time data warehouse for a wide range of real-time data processing scenarios.

Streaming query pushdown

Fluss uses the Apache Arrow columnar format at its storage layer. It supports query acceleration features such as column pruning, partition pushdown, predicate pushdown, and aggregate pushdown. This significantly reduces I/O overhead and can improve query performance by up to 10 times.

CDC log subscription and full/incremental consistency

Fluss has a built-in Change Data Capture (CDC) mechanism that supports real-time subscription to binlog data, ensuring consistent reads of both full and incremental data. In a real-time data warehouse, this enables deduplication and updates, tiered data governance, and end-to-end low-latency data processing.

Partial column update

Fluss supports partial column updates at the primary key level, enabling real-time merging of multiple data streams and avoiding the large state issues common in traditional dual-stream joins. Even in partial update scenarios, a complete binlog is generated to ensure an end-to-end real-time data pipeline.

Delta Join

Built on the streaming read and indexed point query capabilities of Fluss, Delta Join can replace traditional dual-stream joins, significantly reducing memory and CPU usage. In production environments, this has been shown to reduce Flink memory and CPU consumption by over 86%, while shortening checkpoint time from 90s to 1s.

Secondary index and Key-Value point query

Fluss supports the creation of a secondary index to provide efficient Key-Value point query capabilities. This allows users to directly query and verify data in the DWD/DWS layers without exporting it to another system.

Tiered storage for hot and cold data and lake-stream integration

Fluss supports automatic tiered storage for hot and cold data, balancing performance and cost. It also integrates seamlessly with lake formats like Apache Paimon to enable unified management and sharing of stream and lake data, improving the cost-effectiveness of your real-time data warehouse.

Product architecture

image

1. Lake-stream integration architecture

Features
  • Uses a data lake as the underlying storage for low-cost persistence of historical data, while leveraging the lake ecosystem for efficient batch queries and analytics.

  • Includes a built-in streaming-to-lake capability, allowing stream data to be written in real time and become immediately visible without relying on external data integration tools.

  • Supports intelligent tiered management of data between stream storage and the data lake. When data is queried or consumed, the system automatically merges the tiers to provide a unified access view.

  • Adopts an open-source architecture compatible with major lake formats (such as DLF, Paimon, Iceberg, and Lance) and query engines (such as Flink, Spark, and StarRocks), allowing for flexible expansion.

Advantages
  • Unified stream and batch processing: Supports both stream and batch processing within a single framework, reducing architectural complexity and operational costs.

  • Unified data storage: Streams and the lake share a single copy of data, avoiding redundant storage and significantly reducing overall storage costs.

  • Unified data governance: Provides unified access control and metadata management to simplify data governance, ensuring data security and consistency.

2. Distributed cluster architecture

Structure
  • Each Fluss cluster consists of one Coordinator and multiple TabletServer nodes.

  • The Coordinator is responsible for metadata management, task scheduling, and global coordination.

  • TabletServer nodes are responsible for the actual data storage, computation, and query execution.

Advantages
  • High availability: A multi-replica backup mechanism and distributed node design ensure high data reliability and availability.

  • Scalability: Supports on-demand scaling by adding TabletServer nodes or clusters to meet growing business needs.

  • A cold storage layer archives data long-term, reducing the storage cost of hot data.

3. Feature support

Core features
  • Lake-stream integration

    Achieves deep integration of data lakes and stream processing, supporting both real-time data processing and historical data analysis.

  • Streaming query pushdown

    Pushes query logic, such as column pruning and partition pushdown, down to the TabletServer layer. This reduces data transfer and improves streaming query performance.

  • Delta Join

    Supports efficient join analysis between streaming and static data, which is suitable for use cases like real-time recommendations and user profiling.

  • Real-time updates and point queries

    Supports real-time updates and partial column updates at millions of QPS, as well as dimension table point queries. It also generates real-time CDC change logs for seamless integration with Flink to build an end-to-end real-time data warehouse.

Basic features
  • Monitoring and alerting

    Monitors system status in real time and provides alerts to ensure prompt issue detection and resolution.

  • Database and table management

    Supports operations such as creating, modifying, and deleting databases and tables for easy data management and access control.

  • Intra-city disaster recovery

    Provides intra-city disaster recovery to ensure data security and business continuity.

  • Automatic upgrade

    Supports online upgrades to reduce downtime and ensure continuous service availability.

Benefits

Unified stream and batch lakehouse

image

Fluss is designed to help organizations address the common challenges of real-time data processing and extract more value from data. The key benefits include:

  • Real-time performance and low latency

    Fluss supports millisecond-level reads and writes. You can combine it with stream computing engines like Apache Flink to build a high-throughput, low-latency real-time data warehouse that meets the strict timeliness requirements of use cases such as financial trading and e-commerce recommendations.

  • Cost optimization

    The tiered storage strategy for hot and cold data significantly reduces storage costs. Meanwhile, the lake-stream integration architecture minimizes data redundancy and redundant computations, optimizing the total cost of ownership (TCO).

  • Improved data analytics efficiency

    Technologies like query pushdown and Delta Join dramatically improve query performance, shorten the data processing pipeline, and accelerate business insights.

  • Data governance and security

    A unified data management platform supports multi-tenant isolation, access control, and operation auditing to ensure data security and compliance.

  • Flexible scalability

    The service supports on-demand elastic scaling, allowing you to easily handle business peaks or sudden traffic surges without worrying about resource bottlenecks.