A match query performs an approximate match on data in a table. Tablestore uses the configured tokenizer to split the values in Text columns and the search query into tokens, and then searches for these tokens. To perform high-performance fuzzy queries on Text columns that use fuzzy tokenization, use a match phrase query.
Scenarios
A match query finds data that contains a specific phrase. When used with tokenization, a match query enables full-text indexing. This is useful in scenarios such as data analytics, content search, knowledge management, social media analysis, log analysis, AI chat systems, and compliance reviews. For example, you can quickly filter products on an E-commerce platform whose titles, descriptions, or labels contain keywords entered by a user. You can also quickly locate error messages or abnormal operations in logs.
Function overview
A match query performs an approximate match on data in a table. For example, consider a row with a `title` column of the Text type that has the value "Hángzhōu Xīhú Fēngjǐngqū". If this column uses single-word tokenization and the search query is "hú fēng", the row is matched.
When you use a match query, you must specify the column to match and the search query. A row meets the query condition if any of the tokens from the search query exist in the specified column.
You can also configure parameters such as the minimum number of matches, query weight, columns to return, whether to return the total number of matched rows, and the sorting method for the result set.
API operations
The API operations for a match query are Search or ParallelScan. The specific query type is MatchQuery.
Parameters
|
Parameter |
Description |
|
fieldName |
The column to match. A match query can be applied to Text type columns. |
|
text |
The search query, which is the value to match. For Text columns, the search query is split into multiple tokens based on the tokenizer type set when you created the search index. If no tokenizer was set, single-word tokenization is used by default. For example, for a Text column that uses single-word tokenization, the search query "this is" can match "..., this is tablestore", "is this tablestore", "tablestore is cool", "this", and "is". |
|
query |
Set the query type to matchQuery. |
|
offset |
The starting position for the query. |
|
limit |
The maximum number of results to return for the query. To get only the row count without any data, set limit to 0. This does not return any rows. |
|
minimumShouldMatch |
The minimum number of matches. A row is returned only if the value in the `fieldName` column contains at least the minimum number of matching tokens. Note
`minimumShouldMatch` must be used with the OR logical operator. |
|
operator |
The logical operator. The default is OR, which means a row meets the query condition if some of the tokens match. If you set `operator` to AND, a row meets the query condition only if all tokens are present in the column value. |
|
getTotalCount |
Specifies whether to return the total number of matched rows. The default is false, which means the total count is not returned. Returning the total number of matched rows can affect query performance. |
|
weight |
The query weight. This is used for score-based sorting in full-text index scenarios. This parameter specifies the weight for score calculation on a column. A larger value results in a higher score. The value must be a positive floating-point number. This parameter does not affect the number of results returned, only their scores. |
|
tableName |
The name of the data table. |
|
indexName |
The name of the search index. |
|
columnsToGet |
Specifies whether to return all columns. This includes the `returnAll` and `columns` settings. `returnAll` is false by default, which means not all columns are returned. In this case, you can use `columns` to specify which columns to return. If you do not specify columns, only the primary key columns are returned. If you set `returnAll` to true, all columns are returned. |
Notes
Search Index provides only basic BM25 relevance scoring and does not support custom relevance models.
Usage
You can perform a match query using the console, the command line interface, or a software development kit (SDK). Before you begin, complete the following preparations.
Use an Alibaba Cloud account or a RAM user with the required permissions for Table Store operations. To grant permissions to a RAM user, see Grant permissions to a RAM user by using a RAM policy.
If you use an SDK or a command-line tool, create an AccessKey for your Alibaba Cloud account or RAM user if you do not have one.
You have created a data table.
A Search Index has been created for the data table.
If you use an SDK, initialize the Tablestore Client.
If you use the command-line tool, download and start the tool, then configure the connection to your instance and select the target table. For more information, see Download the command-line tool, Start the tool and configure connection information, and Data table operations.
Billing
Querying data by using a Search Index consumes read throughput. For more information, see Search Index metering and billing.
FAQ
References
Search Index supports various query types for multi-dimensional data queries, including term query, terms query, match all query, match query, phrase match query, range query, prefix query, suffix query, wildcard query, token-based wildcard query, boolean query, geo query, nested query, vector search, and exists query.
When you query data, you can sort and paginate the result set or perform collapsing (deduplication).
For data analysis, such as finding the maximum or minimum value, calculating a sum, or counting rows, you can use the statistical aggregation or SQL query features.
To quickly export data regardless of the result set order, you can use the Parallel Scan feature.