Alibaba Cloud Fluss for Apache Fluss is a high-performance, columnar stream storage system built on Apache Fluss. It delivers millisecond-level read and write responses, real-time data updates, and partial column updates. Fluss also provides a Change Data Capture (CDC) log subscription feature, which ensures consistent reads of both full and incremental data for real-time data consumption scenarios.
Prerequisites
Before you begin, it helps to understand the following concepts:
Product overview
The lake-stream integration architecture of Alibaba Cloud Fluss for Apache Fluss enables seamless integration with data lake formats like Apache Paimon. This approach not only ensures high data timeliness but also significantly reduces the cost of building and maintaining a real-time data warehouse. Through its unified data flow and lake-stream collaboration, Fluss helps organizations break down data silos and accelerate value discovery and efficient data sharing within the data lake.
|
Features |
Description |
|
Millisecond-level stream reads and writes |
Fluss supports low-latency streaming read and write operations. You can use it with Apache Flink to build a high-throughput, low-latency real-time data warehouse for a wide range of real-time data processing scenarios. |
|
Streaming query pushdown |
Fluss uses the Apache Arrow columnar format at its storage layer. It supports query acceleration features such as column pruning, partition pushdown, predicate pushdown, and aggregate pushdown. This significantly reduces I/O overhead and can improve query performance by up to 10 times. |
|
CDC log subscription and full/incremental consistency |
Fluss has a built-in Change Data Capture (CDC) mechanism that supports real-time subscription to binlog data, ensuring consistent reads of both full and incremental data. In a real-time data warehouse, this enables deduplication and updates, tiered data governance, and end-to-end low-latency data processing. |
|
Partial column update |
Fluss supports partial column updates at the primary key level, enabling real-time merging of multiple data streams and avoiding the large state issues common in traditional dual-stream joins. Even in partial update scenarios, a complete binlog is generated to ensure an end-to-end real-time data pipeline. |
|
Delta Join |
Built on the streaming read and indexed point query capabilities of Fluss, Delta Join can replace traditional dual-stream joins, significantly reducing memory and CPU usage. In production environments, this has been shown to reduce Flink memory and CPU consumption by over 86%, while shortening checkpoint time from 90s to 1s. |
|
Secondary index and Key-Value point query |
Fluss supports the creation of a secondary index to provide efficient Key-Value point query capabilities. This allows users to directly query and verify data in the DWD/DWS layers without exporting it to another system. |
|
Tiered storage for hot and cold data and lake-stream integration |
Fluss supports automatic tiered storage for hot and cold data, balancing performance and cost. It also integrates seamlessly with lake formats like Apache Paimon to enable unified management and sharing of stream and lake data, improving the cost-effectiveness of your real-time data warehouse. |
Product architecture

1. Lake-stream integration architecture
Features
-
Uses a data lake as the underlying storage for low-cost persistence of historical data, while leveraging the lake ecosystem for efficient batch queries and analytics.
-
Includes a built-in streaming-to-lake capability, allowing stream data to be written in real time and become immediately visible without relying on external data integration tools.
-
Supports intelligent tiered management of data between stream storage and the data lake. When data is queried or consumed, the system automatically merges the tiers to provide a unified access view.
-
Adopts an open-source architecture compatible with major lake formats (such as DLF, Paimon, Iceberg, and Lance) and query engines (such as Flink, Spark, and StarRocks), allowing for flexible expansion.
Advantages
-
Unified stream and batch processing: Supports both stream and batch processing within a single framework, reducing architectural complexity and operational costs.
-
Unified data storage: Streams and the lake share a single copy of data, avoiding redundant storage and significantly reducing overall storage costs.
-
Unified data governance: Provides unified access control and metadata management to simplify data governance, ensuring data security and consistency.
2. Distributed cluster architecture
Structure
-
Each Fluss cluster consists of one Coordinator and multiple TabletServer nodes.
-
The Coordinator is responsible for metadata management, task scheduling, and global coordination.
-
TabletServer nodes are responsible for the actual data storage, computation, and query execution.
Advantages
-
High availability: A multi-replica backup mechanism and distributed node design ensure high data reliability and availability.
-
Scalability: Supports on-demand scaling by adding TabletServer nodes or clusters to meet growing business needs.
-
A cold storage layer archives data long-term, reducing the storage cost of hot data.
3. Feature support
Core features
-
Lake-stream integration
Achieves deep integration of data lakes and stream processing, supporting both real-time data processing and historical data analysis.
-
Streaming query pushdown
Pushes query logic, such as column pruning and partition pushdown, down to the TabletServer layer. This reduces data transfer and improves streaming query performance.
-
Delta Join
Supports efficient join analysis between streaming and static data, which is suitable for use cases like real-time recommendations and user profiling.
-
Real-time updates and point queries
Supports real-time updates and partial column updates at millions of QPS, as well as dimension table point queries. It also generates real-time CDC change logs for seamless integration with Flink to build an end-to-end real-time data warehouse.
Basic features
-
Monitoring and alerting
Monitors system status in real time and provides alerts to ensure prompt issue detection and resolution.
-
Database and table management
Supports operations such as creating, modifying, and deleting databases and tables for easy data management and access control.
-
Intra-city disaster recovery
Provides intra-city disaster recovery to ensure data security and business continuity.
-
Automatic upgrade
Supports online upgrades to reduce downtime and ensure continuous service availability.
Benefits
Unified stream and batch lakehouse

Fluss is designed to help organizations address the common challenges of real-time data processing and extract more value from data. The key benefits include:
-
Real-time performance and low latency
Fluss supports millisecond-level reads and writes. You can combine it with stream computing engines like Apache Flink to build a high-throughput, low-latency real-time data warehouse that meets the strict timeliness requirements of use cases such as financial trading and e-commerce recommendations.
-
Cost optimization
The tiered storage strategy for hot and cold data significantly reduces storage costs. Meanwhile, the lake-stream integration architecture minimizes data redundancy and redundant computations, optimizing the total cost of ownership (TCO).
-
Improved data analytics efficiency
Technologies like query pushdown and Delta Join dramatically improve query performance, shorten the data processing pipeline, and accelerate business insights.
-
Data governance and security
A unified data management platform supports multi-tenant isolation, access control, and operation auditing to ensure data security and compliance.
-
Flexible scalability
The service supports on-demand elastic scaling, allowing you to easily handle business peaks or sudden traffic surges without worrying about resource bottlenecks.