Lake-stream integration is a new data architecture that merges data lakes with real-time streams.
Lake-stream integration

Pain points of traditional lakehouses and stream processing
-
Storage silos: A single storage layer cannot meet the requirements for both real-time performance and low cost.
-
Complex architecture: The Lambda architecture requires developers to maintain two completely separate systems.
-
Storage redundancy: Data is stored redundantly across different systems, which consumes both storage and compute resources.
-
Fragmented queries: Without a unified data view, different compute engines and processing logic can easily lead to ambiguous data semantics.
Core concepts and value of lake-stream integration
Lake-stream integration is a new data architecture that merges data lakes with real-time streams. It aims to break down the boundaries between traditional batch and stream processing. This integration enables unified real-time and batch data processing with a single source for storage, metadata, and access. This architecture is often used to solve data timeliness issues in data lakes, such as reducing latency from minutes to milliseconds. It also provides Online Analytical Processing (OLAP) capabilities for data streams.
This architecture deeply integrates the real-time stream storage (Fluss) with a Paimon-based data lake, with metadata managed by DLF. This integration creates 'one copy of data, two views'. A single logical table provides both low-latency access to real-time data and high-throughput analytics for historical data.
Core principles
-
Automatic synchronization: When this feature is enabled on a Fluss cluster, the system automatically creates a Flink synchronization job. This job continuously writes the real-time event stream from Fluss to a Paimon table.
-
Union read: When a query is run on a table, the Flink query engine automatically merges real-time data from Fluss with historical data from Paimon. It then deduplicates the data by primary key and presents a single, consistent, and up-to-date table.
-
Unified metadata: DLF acts as a unified metadata service. It manages Paimon table information, such as table schemas, partitions, and permissions. This ensures that the data can be seamlessly accessed by multiple compute engines, such as Flink, Spark, and Presto.
-
Data sharing: Fluss and Paimon work together to create an integrated stream-batch storage architecture. Fluss serves as the real-time data layer, providing millisecond-latency writes and access for hot data. Paimon serves as the historical data layer and handles long-term, cost-effective storage for historical data. Data is shared between both layers to prevent redundant storage.
Key capabilities
-
Reduced architectural complexity: Eliminates the dual development paths, duplicate storage, and duplicated Operations and Maintenance (O&M) costs associated with traditional separate stream and batch systems.
-
Guaranteed data consistency: Ensures consistent results for both real-time and batch analytics using a unified data model and compute engine.
-
Simplified O&M: Provides an end-to-end monitoring dashboard that shows metrics such as synchronization latency, job status, and health scores. It also integrates O&M functions such as one-click start/stop and alert settings.
-
Increased agility: Developers can focus on business logic without worrying about the physical location of the underlying data. This allows them to quickly build applications for scenarios such as real-time data warehouses, real-time reports, and AI feature engineering.