Full-text search
Full-text search is an information retrieval technology that lets you quickly and accurately find information relevant to a user's query from large amounts of text data.PolarDB for PostgreSQLProvides a set of full-text search capabilities.
Background
Unlike traditional keyword search, full-text search finds information relevant to a user's query quickly and accurately from massive amounts of text data. It processes the entire content of a document, not just specific fields or tags.
A full-text search system typically includes the following steps:
-
Text preprocessing: This step tokenizes text, removes stop words, and performs stemming to improve search efficiency and accuracy.
-
Index creation: An index is created for the processed text. This process typically uses an inverted index structure to record the position of each word and the documents in which it appears.
-
Query processing: When a user enters a query, the system analyzes the query statement and converts it into a suitable format for searching.
-
Result sorting: The search results are sorted based on a relevance algorithm to return the results that best match the user's query.
Scenarios
Full-text search is used in a wide range of scenarios. Some typical scenarios include the following:
-
Document management systems: You can quickly find internal company documents, reports, and clauses to improve work efficiency.
-
Online search engines: Full-text search technology helps users quickly find the information they need.
-
Academic research: In academic databases and literature management tools, researchers can use full-text search to quickly locate relevant literature and data.
-
E-commerce: E-commerce platforms use full-text search to help customers quickly find the products they need and improve the shopping experience.
-
Social media: Users can search for posts, comments, and images by keyword, which makes it easier to retrieve information.
-
Legal document retrieval: Lawyers and legal professionals use full-text search tools to quickly find relevant cases and legal clauses.
-
Medical record management: Hospitals can use full-text search technology to quickly find patient information and history in medical records and reports.
-
Customer support: Online customer service systems use full-text search to help customers quickly find answers to frequently asked questions (FAQs) and support documents.
-
Content management systems: Websites and blogs use full-text search to help visitors quickly find relevant articles and materials.
-
Libraries and information retrieval: Library information retrieval systems use full-text search technology to make it easy for readers to find books and articles.
Feature overview
Tokenization
PolarDB for PostgreSQLFull-text indexes allow preprocessing of documents and saving an index for subsequent fast searches. Preprocessing includes:
-
Parsing documents into tokens. This step identifies different types of tokens, such as numbers, words, compound words, and email addresses, so that they can be processed in different ways.
-
Converting tokens into lexemes. A lexeme is a normalized string, similar to a token, that unifies different forms of the same word into a single representation. For example, normalization usually includes converting uppercase letters to lowercase and often involves removing suffixes, such as 's' or 'es' in English. This process allows searches to recognize different variations of the same word and eliminates the need to enter all possible variants.
-
Storing preprocessed documents for search optimization. For example, each document can be represented as an ordered array of normalized lexemes. In addition to lexemes, position information is usually stored for proximity ranking. As a result, a document with a dense area of query words ranks higher than a document with scattered query words.
tsvector
tsvector is a data type provided byPolarDB for PostgreSQLA data type for full-text search. This data type can efficiently store processed text for fast search and matching. tsvector is used to store the lexical index of a document, usually consisting of vocabulary and their positions in the text.
pg_bigm
pg_bigmisPolarDB for PostgreSQLAn extension module for fuzzy search, especially for approximate string matching. This extension is most commonly used in applications with large text data, such as search engines and content management systems. The core idea is to improve text search efficiency and accuracy through n-grams.
The pg_bigm extension is effective for prefix and suffix fuzzy queries, such as like '%xxxx%'.
pg_trgm
pg_trgm is an extension forPolarDB for PostgreSQLAn extension module that provides support for trigrams. A trigram is a method for string similarity search, especially suitable for fuzzy matching and text similarity queries. pg_trgm mainly improves query efficiency for large text data sets by creating indexes and query operators, and is commonly used in full-text search, auto-completion, spell correction, and other scenarios.
The pg_trgm extension is effective for prefix and suffix fuzzy queries, such as like '%xxxx%'.
Chinese tokenization
In the Chinese language, words are the smallest morphemic units. Unlike English, Chinese is written without spaces between words. Because of this characteristic, it is difficult to obtain tokenization results that match Chinese semantics when you use the default full-text search engine in PostgreSQL.
For Chinese tokenization,PolarDB for PostgreSQLSupports two full-text search capabilities: pg_jieba and Zhparser.
pg_jieba
Jieba is one of the most popular Chinese tokenization libraries. It can accurately identify and tokenize words in Chinese sentences. Thepg_jiebaThe extension introduces Jieba tokenization capabilities into the database, enabling efficient Chinese text tokenization to support full-text search.
Zhparser
Simple Chinese Word Segmentation (SCWS) is an open source Chinese tokenization engine based on a word frequency dictionary. It can accurately segment a block of Chinese text into words.
ZhparserA Chinese tokenization extension based on SCWS capabilities. While compatible with existing PostgreSQL full-text search capabilities, it provides rich configuration options and supports custom dictionaries.
Index
PolarDB for PostgreSQLSupports multiple index structures for full-text search.
GIN index
Generalized Inverted Index (GIN) is an index type in PostgreSQL that supports full-text search. Using a GIN index, you can efficiently perform full-text searches, especially when you handle large-scale text data. A GIN index allows for fast query operations, especially when you process complex text queries that use tsvector and tsquery. GIN indexes also support other data types, such as JSONB.
RUM index
RUMIs a PostgreSQL extension that provides a RUM index type for full-text search and other indexing needs. This index is designed to improve full-text search performance, especially in scenarios where documents need to be ranked.
A RUM index is an inverted index similar to the built-in Generalized Inverted Index (GIN). The main difference is that a RUM index can store additional information. This allows for faster results when you perform sorting or other operations. For example, in a full-text search, a RUM index can store the position of words in a document. You can then use this position information to calculate relevance rankings during a query.
Query processing
tsquery
tsquery is a feature for full-text search that is specifically designed for querying text data. It allows users to create complex search conditions to find information quickly and effectively in large-scale text data.PolarDB for PostgreSQLAlso provides the to_tsquery method to convert text to tsquery. Combined with tsvector and full-text search operators, full-text search queries can be completed.
tsquery supports the@@(Containment) operators and Boolean operators&( AND)、|(OR) and!(NOT), which makes it easy to construct combined condition search queries.
Ranking
ts_rank
ts_rank is a PostgreSQL function used for full-text search. It is mainly used to calculate a relevance score between a document and a query. This score can be used to evaluate the importance or relevance of a document for a specific query.