User guide for sparse vector indexes
Hologres V5.0 and later support the SPARSEVECTOR data type and sparse vector indexes built with the SINDI (Sparse Inverted No-redundancy Distance Index) algorithm. A sparse vector index combines an inverted structure with SIMD acceleration to deliver efficient approximate nearest neighbor (ANN) search over high-dimensional sparse vectors, such as those used in text retrieval and recommendation systems. This topic describes when to use sparse vector indexes, how to manage them, how to run searches, and how to tune performance.
Comparison with HGraph indexes
Hologres provides two types of vector index: sparse vector indexes (SINDI) and HGraph indexes. Each type targets a different data type and a different class of workload. Pick the one that matches your data.
Sparse vector indexes (SINDI)
A sparse vector stores only its non-zero elements and their positions, in {index:value} form. For example, a one-million-dimensional vector with only 100 non-zero elements needs just 100 index-value pairs, which is far cheaper to store and compute than a dense representation.
Sparse vector indexes fit the following workloads:
-
Text retrieval: Models such as SPLADE and BGE-M3 produce text embeddings that are typically high-dimensional sparse vectors.
-
Recommendation systems: Sparse vectors built from user-item interaction features.
-
Advertising features: Sparse feature vectors built from user profiles and ad attributes.
HGraph indexes
HGraph is the dense vector index in Hologres. It uses a graph-based algorithm and targets ANN search over low-dimensional dense vectors. For more information, see Work with HGraph for vector search.
Choose an index type
|
Dimension |
Sparse vector index (SINDI) |
HGraph index |
|
Data type |
SPARSEVECTOR |
float4[] |
|
Vector dimensions |
High-dimensional and sparse (up to tens of millions of dimensions, with few non-zero elements) |
Low-dimensional and dense (hundreds to thousands of dimensions) |
|
Index algorithm |
SINDI |
HGraph |
|
Supported distance methods |
InnerProduct only |
Euclidean, InnerProduct, Cosine |
|
Storage mode |
Full in-memory |
Full in-memory, or hybrid memory-disk |
|
Hybrid search |
Supported (combine sparse and dense vector queries in the same SQL statement) |
Supported |
Recommendations:
-
Choose a sparse vector index if your data comes from sparse encoders such as SPLADE or BGE-M3, or if your vectors have extremely high dimensions but mostly zero elements, as with text embeddings and recommendation features.
-
Choose an HGraph index if your data consists of low-dimensional dense vectors such as image embeddings or semantic vectors, or if you need Cosine distance, Euclidean distance, or a disk-based storage mode.
Limits
-
Sparse vector indexes are supported only in Hologres V5.0 and later.
-
Only column-oriented tables and hybrid row-column tables support sparse vector indexes. Row-oriented tables do not.
-
Data in Mem Tables has no sparse vector index. A search request falls back to brute-force computation for that portion of the data.
-
A sparse vector column cannot have both a sparse vector index and an HGraph index. Each column supports only one type of vector index.
-
Only the full in-memory mode is supported. Watch memory usage for large datasets, and size your instance specification based on your data volume.
-
When you create a sparse vector index,
algorithmin thevectorsproperty supports onlySindi, anddistance_methodsupports onlyInnerProduct. Theproxima_vectorsproperty is not supported.
Manage sparse vector indexes
Create an index
Create a SINDI index on a sparse vector column either when you create the table or after the table exists.
Syntax:
-- Create the index when you create the table
CREATE TABLE <TABLE_NAME> (
<COLUMN_NAME> SPARSEVECTOR,
...
) WITH (
vectors = '{
"<COLUMN_NAME>": {
"algorithm": "Sindi",
"distance_method": "InnerProduct",
"builder_params": {
"use_reorder": <TRUE|FALSE>,
"term_id_limit": <INTEGER>,
"doc_prune_ratio": <FLOAT>,
"window_size": <INTEGER>
}
}}'
);
-- Create the index after the table exists
ALTER TABLE <TABLE_NAME> SET (
vectors = '{
"<COLUMN_NAME>": {
"algorithm": "Sindi",
"distance_method": "InnerProduct",
"builder_params": {
"use_reorder": <TRUE|FALSE>,
"term_id_limit": <INTEGER>,
"doc_prune_ratio": <FLOAT>,
"window_size": <INTEGER>
}
}}'
);
Data type:
Sparse vector columns use the SPARSEVECTOR data type. A SPARSEVECTOR value is an unordered set of index-value pairs written in the text format '{index:value, index:value, ...}', where index positions are integers that start at 1 and must be smaller than the term_id_limit value, and values are floating-point numbers.
Example: '{1:0.1, 50:0.8, 200:0.3}'
Parameters:
|
Parameter |
Required |
Default value |
Description |
|
algorithm |
Yes |
None |
The index algorithm. Only Sindi is supported. |
|
distance_method |
Yes |
None |
The distance calculation method. Only InnerProduct is supported. |
|
use_reorder |
No |
false |
Whether to enable reordering. Reordering improves the recall rate. Set it to true for high-precision workloads. |
|
term_id_limit |
No |
1000000 |
The upper limit on the feature dimensions of the sparse vector index. Index positions in the text format (1-based) must be smaller than this value. Set it according to the largest index position in your data. |
|
doc_prune_ratio |
No |
0.0 |
The document pruning ratio applied when the index is built. Valid values: [0.0, 1.0). A larger value improves query performance but lowers the recall rate. |
|
window_size |
No |
50000 |
The window size used to build the index. Valid values: [10000, 60000]. This value affects memory usage and build performance. |
|
use_quantization |
No |
false |
Whether to enable quantization. Quantization reduces memory usage but may lower the recall rate. |
The following example creates the index together with the table:
CREATE TABLE sparse_test (
id INT PRIMARY KEY,
vec SPARSEVECTOR
) WITH (
vectors = '{
"vec": {
"algorithm": "Sindi",
"distance_method": "InnerProduct",
"builder_params": {
"use_reorder": true,
"term_id_limit": 1000000,
"doc_prune_ratio": 0.4,
"window_size": 50000
}
}}'
);
Modify an index
Use ALTER TABLE to change the build parameters of a sparse vector index.
-- Change the document pruning ratio
ALTER TABLE sparse_test SET (
vectors = '{
"vec": {
"algorithm": "Sindi",
"distance_method": "InnerProduct",
"builder_params": {
"use_reorder": true,
"term_id_limit": 1000000,
"doc_prune_ratio": 0.6,
"window_size": 50000
}
}}'
);
New parameters take effect for indexes that are built after the change. To apply them sooner, trigger compaction manually:
VACUUM <SCHEMA_NAME>.<TABLE_NAME>;
For more information about compaction, see Compact data files.
Drop an index
Set vectors to an empty JSON object to drop the sparse vector index.
ALTER TABLE sparse_test SET (vectors = '{}');
View an index
Query the hologres.hg_table_properties system table to view the configuration of a sparse vector index.
SELECT property_value
FROM hologres.hg_table_properties
WHERE table_name = 'sparse_test'
AND property_key = 'vectors';
Search with sparse vectors
Hologres provides the following distance functions for sparse vector search.
|
Function |
Description |
|
approx_inner_product_distance(sparsevector, sparsevector) |
Approximate search. Uses the SINDI index for acceleration and returns the inner product distance. |
|
inner_product_distance(sparsevector, sparsevector) |
Exact search with brute-force computation. Does not use the index and returns the inner product distance. |
Sparse vectors support only the InnerProduct distance method. Functions such as approx_cosine_distance and approx_euclidean_distance are not supported. To work with cosine distance, normalize your vectors first and then use InnerProduct.
Approximate search
Use approx_inner_product_distance with ORDER BY ... DESC LIMIT k to run a top-k approximate search. The query uses the SINDI index automatically.
SELECT id, approx_inner_product_distance(vec, '{1:0.1, 2:0.2, 3:0.3}') AS dist
FROM sparse_test
ORDER BY dist DESC
LIMIT 5;
Exact search
Use inner_product_distance to run an exact search with brute-force computation. This function bypasses the index, so it suits small datasets and cases that require exact results.
SELECT id, inner_product_distance(vec, '{1:0.1, 2:0.2, 3:0.3}') AS dist
FROM sparse_test
ORDER BY dist DESC
LIMIT 5;
Verify that the index is used
Run EXPLAIN ANALYZE to inspect the execution plan and confirm whether the query uses the sparse vector index.
EXPLAIN ANALYZE
SELECT product_id,
approx_inner_product_distance(feature, '{1:0.1, 50:0.7}') AS dist
FROM product_features
ORDER BY dist DESC
LIMIT 5;
If Vector Filter appears for the sparse vector column in the execution plan, the query uses the sparse vector index.
Vector Filter: VectorCond => KNN: '5'::bigint distance_method: approx_inner_product_distance
When you use inner_product_distance, the execution plan shows Seq Scan without Vector Filter. This means the query does not use the index and runs brute-force computation instead.
Examples
The following example walks through the complete workflow for a sparse vector index: create the table, write data, and run a query.
-- 1. Create a table with a sparse vector index
CREATE TABLE product_features (
product_id INT PRIMARY KEY,
feature SPARSEVECTOR,
product_name TEXT
) WITH (
vectors = '{
"feature": {
"algorithm": "Sindi",
"distance_method": "InnerProduct",
"builder_params": {
"use_reorder": true,
"term_id_limit": 100000,
"window_size": 50000
}
}}'
);
-- 2. Write data. Sparse vectors use the text format {index:value, ...}, where index starts at 1.
INSERT INTO product_features VALUES
(1, '{1:0.1, 50:0.8, 200:0.3}', 'Product A'),
(2, '{1:0.2, 100:0.6}', 'Product B'),
(3, '{50:0.9, 200:0.1, 500:0.4}', 'Product C'),
(4, '{1:0.1, 100:0.5, 200:0.7}', 'Product D'),
(5, '{50:0.3, 500:0.8}', 'Product E');
-- 3. Query the top 5 products most similar to the target vector
SELECT product_id, product_name,
approx_inner_product_distance(feature, '{1:0.1, 50:0.7, 200:0.5}') AS similarity
FROM product_features
ORDER BY similarity DESC
LIMIT 5;
Sample result:
|
product_id |
product_name |
similarity |
|
1 |
Product A |
0.72 |
|
3 |
Product C |
0.68 |
|
4 |
Product D |
0.36 |
|
5 |
Product E |
0.21 |
|
2 |
Product B |
0.02 |
The preceding values are examples. Your actual results depend on your data.
Performance tuning
Tune build parameters
doc_prune_ratio is the build parameter with the largest impact on query performance. Increasing it speeds up queries noticeably, at the cost of some recall.
|
doc_prune_ratio value |
Effect |
|
0.0 |
No pruning. Highest recall rate—recall@10 reaches 99.99% with default parameters—but lower query performance. |
|
0.4 (recommended) |
A good balance between performance and recall. Query performance improves about 10 times compared with 0.0, with only a slight drop in recall. |
|
Greater than 0.4 |
Query performance improves further, but the recall rate drops more noticeably. |
Setting use_reorder to true improves query performance while keeping the recall rate high. Enable it for high-precision workloads.
Tune query parameters
Set the following GUC parameters to balance performance and recall at query time.
|
GUC parameter |
Default value |
Description |
|
hg_vector_sindi_n_candidate |
0 |
The candidate set size at query time. The default value 0 means topk × 500. A larger value increases the recall rate but also increases query latency. |
|
hg_vector_sindi_query_prune_ratio |
0.0 |
The pruning ratio at query time. Valid values: [0.0, 0.9]. A larger value speeds up queries but lowers the recall rate. |
|
hg_vector_sindi_term_prune_ratio |
0.0 |
The term pruning ratio at query time. Valid values: [0.0, 0.9]. Not recommended, because it has a large impact on the recall rate. |
|
hg_vector_sindi_use_term_lists_heap_insert |
true |
Whether to use heap insertion optimization. |
Example:
-- Set the query parameters
SET hg_vector_sindi_n_candidate = 500;
SET hg_vector_sindi_query_prune_ratio = 0.8;
-- Run the query
SELECT id, approx_inner_product_distance(vec, '{1:0.1, 2:0.2}') AS dist
FROM sparse_test
ORDER BY dist DESC
LIMIT 10;
-- Reset to the default values
RESET hg_vector_sindi_n_candidate;
RESET hg_vector_sindi_query_prune_ratio;
Recommended configurations by scenario
The following recommendations are measured on the sparse-full dataset, which contains 8.84 million vectors with 30,109 dimensions and 99.6% sparsity.
|
Scenario |
doc_prune_ratio (build) |
query_prune_ratio (query) |
n_candidate (query) |
Expected recall rate |
Expected P99 latency |
|
High precision (RAG text retrieval) |
0.0–0.2 |
0.7 |
500–1000 |
Greater than 99.8% |
About 15 ms |
|
Balanced |
0.4 |
0.8 |
300–500 |
Greater than 99.4% |
About 13 ms |
|
High performance (real-time search) |
0.4–0.6 |
0.9 |
500–1000 |
Greater than 96% |
About 10 ms |
FAQ
-
Q: Creating a sparse vector index fails with
For Sindi algorithm, only 'InnerProduct' distance_method is supported.A: The SINDI algorithm supports only the InnerProduct distance method. Change distance_method to InnerProduct. To work with cosine distance, normalize your sparse vectors first and then use InnerProduct.
-
Q: The error
If property_key is 'proxima_vectors', only Graph algorithm is supportedoccurs.A: Sparse vector indexes do not support the proxima_vectors property. Use the vectors property to create the index.
-
Q: Writing a sparse vector fails with an error that an index position is out of range.
A: Index positions in the sparse vector text format (1-based) must be smaller than the term_id_limit value. Check the largest index position in your data and set term_id_limit high enough when you create the index. For example, if the largest index position in your data is 500000, set term_id_limit to a value greater than 500000.
-
Q: Do sparse vector indexes support extra_columns?
A: No. The SINDI algorithm does not currently support the extra_columns parameter. Do not include it in builder_params.
-
Q: How do I tell whether a query uses the sparse vector index?
A: Run EXPLAIN ANALYZE to inspect the execution plan. If the plan contains the Vector Filter keyword, the query uses the sparse vector index. If you use the inner_product_distance function for exact search, the plan shows Seq Scan, which means the index is not used.