Apache Druid is a real-time analytics system that uses distributed memory. It provides fast and interactive queries and analysis for large scale datasets.
Basic features
Apache Druid has the following features:
Sub-second interactive queries, such as multi-dimensional filtering, ad hoc grouping by property, and fast data aggregation.
Real-time data consumption.
Concurrent online queries from multiple tenants.
Processes petabytes of data and hundreds of billions of events, and supports thousands of concurrent queries per second.
High availability (HA) and rolling upgrades.
Scenarios
Real-time data analytics is the most common scenario for Apache Druid. This scenario includes a wide range of applications, such as:
Real-time metrics monitoring
Recommendation models
Advertising platforms
Search models
Apache Druid architecture
Apache Druid has a well-designed architecture. Multiple components work together to handle the entire data flow, from ingestion to indexing, storage, and querying.
The Druid working layer, which handles data indexing and queries, includes the following components:
The Realtime component ingests data in real time.
The Broker component distributes query tasks, aggregates the results, and returns them to the user.
The Historical component stores indexed historical data in deep storage. Deep storage can be a local disk or a distributed file system, such as Hadoop Distributed File System (HDFS).
The Indexing service includes the following two components:
The Overlord component manages and distributes indexing tasks.
The MiddleManager component executes the indexing tasks.
The Druid segments (Druid manifest) management layer includes the following components:
ZooKeeper: Stores the cluster state and acts as the service discovery component. For example, it manages cluster topology information, Overlord leader election, and indexing tasks.
Coordinator: Manages segments. For example, it handles segment downloads, deletions, and load balancing between Historical nodes.
Metastore: Stores segment metadata and manages various persistent or temporary data for the cluster, such as configuration and audit information.
E-MapReduce enhanced Druid
E-MapReduce Druid is a ready-to-use and fully managed service based on Apache Druid. It includes many improvements, such as integration with E-MapReduce and other Alibaba Cloud services, convenient monitoring and O&M features, and an easy-to-use product interface.
E-MapReduce Druid currently supports the following features:
Uses Object Storage Service (OSS) as deep storage.
Uses files in OSS as a data source for batch indexing.
Streams and indexes data from Simple Log Service, which is similar to Kafka. This feature provides high reliability and exactly-once semantics.
Stores metadata in ApsaraDB RDS.
Integrates the Superset tool.
Supports easy scale-out and scale-in. Scale-in applies only to Task nodes.
Provides rich monitoring metrics and alert rules.
Failover.
Provides high security.
Supports HA.