User guide for sparse vector indexes

Updated at:

Hologres V5.0 and later support the SPARSEVECTOR data type and sparse vector indexes built with the SINDI (Sparse Inverted No-redundancy Distance Index) algorithm. A sparse vector index combines an inverted structure with SIMD acceleration to deliver efficient approximate nearest neighbor (ANN) search over high-dimensional sparse vectors, such as those used in text retrieval and recommendation systems. This topic describes when to use sparse vector indexes, how to manage them, how to run searches, and how to tune performance.

Comparison with HGraph indexes

Hologres provides two types of vector index: sparse vector indexes (SINDI) and HGraph indexes. Each type targets a different data type and a different class of workload. Pick the one that matches your data.

Sparse vector indexes (SINDI)

A sparse vector stores only its non-zero elements and their positions, in {index:value} form. For example, a one-million-dimensional vector with only 100 non-zero elements needs just 100 index-value pairs, which is far cheaper to store and compute than a dense representation.

Sparse vector indexes fit the following workloads:

  • Text retrieval: Models such as SPLADE and BGE-M3 produce text embeddings that are typically high-dimensional sparse vectors.

  • Recommendation systems: Sparse vectors built from user-item interaction features.

  • Advertising features: Sparse feature vectors built from user profiles and ad attributes.

HGraph indexes

HGraph is the dense vector index in Hologres. It uses a graph-based algorithm and targets ANN search over low-dimensional dense vectors. For more information, see Work with HGraph for vector search.

Choose an index type

Dimension

Sparse vector index (SINDI)

HGraph index

Data type

SPARSEVECTOR

float4[]

Vector dimensions

High-dimensional and sparse (up to tens of millions of dimensions, with few non-zero elements)

Low-dimensional and dense (hundreds to thousands of dimensions)

Index algorithm

SINDI

HGraph

Supported distance methods

InnerProduct only

Euclidean, InnerProduct, Cosine

Storage mode

Full in-memory

Full in-memory, or hybrid memory-disk

Hybrid search

Supported (combine sparse and dense vector queries in the same SQL statement)

Supported

Recommendations:

  • Choose a sparse vector index if your data comes from sparse encoders such as SPLADE or BGE-M3, or if your vectors have extremely high dimensions but mostly zero elements, as with text embeddings and recommendation features.

  • Choose an HGraph index if your data consists of low-dimensional dense vectors such as image embeddings or semantic vectors, or if you need Cosine distance, Euclidean distance, or a disk-based storage mode.

Limits

  • Sparse vector indexes are supported only in Hologres V5.0 and later.

  • Only column-oriented tables and hybrid row-column tables support sparse vector indexes. Row-oriented tables do not.

  • Data in Mem Tables has no sparse vector index. A search request falls back to brute-force computation for that portion of the data.

  • A sparse vector column cannot have both a sparse vector index and an HGraph index. Each column supports only one type of vector index.

  • Only the full in-memory mode is supported. Watch memory usage for large datasets, and size your instance specification based on your data volume.

  • When you create a sparse vector index, algorithm in the vectors property supports only Sindi, and distance_method supports only InnerProduct. The proxima_vectors property is not supported.

Manage sparse vector indexes

Create an index

Create a SINDI index on a sparse vector column either when you create the table or after the table exists.

Syntax:

-- Create the index when you create the table
CREATE TABLE <TABLE_NAME> (
    <COLUMN_NAME> SPARSEVECTOR,
    ...
) WITH (
    vectors = '{
    "<COLUMN_NAME>": {
        "algorithm": "Sindi",
        "distance_method": "InnerProduct",
        "builder_params": {
            "use_reorder": <TRUE|FALSE>,
            "term_id_limit": <INTEGER>,
            "doc_prune_ratio": <FLOAT>,
            "window_size": <INTEGER>
        }
    }}'
);

-- Create the index after the table exists
ALTER TABLE <TABLE_NAME> SET (
    vectors = '{
    "<COLUMN_NAME>": {
        "algorithm": "Sindi",
        "distance_method": "InnerProduct",
        "builder_params": {
            "use_reorder": <TRUE|FALSE>,
            "term_id_limit": <INTEGER>,
            "doc_prune_ratio": <FLOAT>,
            "window_size": <INTEGER>
        }
    }}'
);

Data type:

Sparse vector columns use the SPARSEVECTOR data type. A SPARSEVECTOR value is an unordered set of index-value pairs written in the text format '{index:value, index:value, ...}', where index positions are integers that start at 1 and must be smaller than the term_id_limit value, and values are floating-point numbers.

Example: '{1:0.1, 50:0.8, 200:0.3}'

Parameters:

Parameter

Required

Default value

Description

algorithm

Yes

None

The index algorithm. Only Sindi is supported.

distance_method

Yes

None

The distance calculation method. Only InnerProduct is supported.

use_reorder

No

false

Whether to enable reordering. Reordering improves the recall rate. Set it to true for high-precision workloads.

term_id_limit

No

1000000

The upper limit on the feature dimensions of the sparse vector index. Index positions in the text format (1-based) must be smaller than this value. Set it according to the largest index position in your data.

doc_prune_ratio

No

0.0

The document pruning ratio applied when the index is built. Valid values: [0.0, 1.0). A larger value improves query performance but lowers the recall rate.

window_size

No

50000

The window size used to build the index. Valid values: [10000, 60000]. This value affects memory usage and build performance.

use_quantization

No

false

Whether to enable quantization. Quantization reduces memory usage but may lower the recall rate.

The following example creates the index together with the table:

CREATE TABLE sparse_test (
    id INT PRIMARY KEY,
    vec SPARSEVECTOR
) WITH (
    vectors = '{
    "vec": {
        "algorithm": "Sindi",
        "distance_method": "InnerProduct",
        "builder_params": {
            "use_reorder": true,
            "term_id_limit": 1000000,
            "doc_prune_ratio": 0.4,
            "window_size": 50000
        }
    }}'
);

Modify an index

Use ALTER TABLE to change the build parameters of a sparse vector index.

-- Change the document pruning ratio
ALTER TABLE sparse_test SET (
    vectors = '{
    "vec": {
        "algorithm": "Sindi",
        "distance_method": "InnerProduct",
        "builder_params": {
            "use_reorder": true,
            "term_id_limit": 1000000,
            "doc_prune_ratio": 0.6,
            "window_size": 50000
        }
    }}'
);

New parameters take effect for indexes that are built after the change. To apply them sooner, trigger compaction manually:

VACUUM <SCHEMA_NAME>.<TABLE_NAME>;

For more information about compaction, see Compact data files.

Drop an index

Set vectors to an empty JSON object to drop the sparse vector index.

ALTER TABLE sparse_test SET (vectors = '{}');

View an index

Query the hologres.hg_table_properties system table to view the configuration of a sparse vector index.

SELECT property_value
FROM hologres.hg_table_properties
WHERE table_name = 'sparse_test'
  AND property_key = 'vectors';

Examples

The following example walks through the complete workflow for a sparse vector index: create the table, write data, and run a query.

-- 1. Create a table with a sparse vector index
CREATE TABLE product_features (
    product_id INT PRIMARY KEY,
    feature SPARSEVECTOR,
    product_name TEXT
) WITH (
    vectors = '{
    "feature": {
        "algorithm": "Sindi",
        "distance_method": "InnerProduct",
        "builder_params": {
            "use_reorder": true,
            "term_id_limit": 100000,
            "window_size": 50000
        }
    }}'
);

-- 2. Write data. Sparse vectors use the text format {index:value, ...}, where index starts at 1.
INSERT INTO product_features VALUES
    (1, '{1:0.1, 50:0.8, 200:0.3}', 'Product A'),
    (2, '{1:0.2, 100:0.6}', 'Product B'),
    (3, '{50:0.9, 200:0.1, 500:0.4}', 'Product C'),
    (4, '{1:0.1, 100:0.5, 200:0.7}', 'Product D'),
    (5, '{50:0.3, 500:0.8}', 'Product E');

-- 3. Query the top 5 products most similar to the target vector
SELECT product_id, product_name,
       approx_inner_product_distance(feature, '{1:0.1, 50:0.7, 200:0.5}') AS similarity
FROM product_features
ORDER BY similarity DESC
LIMIT 5;

Sample result:

product_id

product_name

similarity

1

Product A

0.72

3

Product C

0.68

4

Product D

0.36

5

Product E

0.21

2

Product B

0.02

Note

The preceding values are examples. Your actual results depend on your data.

Performance tuning

Tune build parameters

doc_prune_ratio is the build parameter with the largest impact on query performance. Increasing it speeds up queries noticeably, at the cost of some recall.

doc_prune_ratio value

Effect

0.0

No pruning. Highest recall rate—recall@10 reaches 99.99% with default parameters—but lower query performance.

0.4 (recommended)

A good balance between performance and recall. Query performance improves about 10 times compared with 0.0, with only a slight drop in recall.

Greater than 0.4

Query performance improves further, but the recall rate drops more noticeably.

Setting use_reorder to true improves query performance while keeping the recall rate high. Enable it for high-precision workloads.

Tune query parameters

Set the following GUC parameters to balance performance and recall at query time.

GUC parameter

Default value

Description

hg_vector_sindi_n_candidate

0

The candidate set size at query time. The default value 0 means topk × 500. A larger value increases the recall rate but also increases query latency.

hg_vector_sindi_query_prune_ratio

0.0

The pruning ratio at query time. Valid values: [0.0, 0.9]. A larger value speeds up queries but lowers the recall rate.

hg_vector_sindi_term_prune_ratio

0.0

The term pruning ratio at query time. Valid values: [0.0, 0.9]. Not recommended, because it has a large impact on the recall rate.

hg_vector_sindi_use_term_lists_heap_insert

true

Whether to use heap insertion optimization.

Example:

-- Set the query parameters
SET hg_vector_sindi_n_candidate = 500;
SET hg_vector_sindi_query_prune_ratio = 0.8;

-- Run the query
SELECT id, approx_inner_product_distance(vec, '{1:0.1, 2:0.2}') AS dist
FROM sparse_test
ORDER BY dist DESC
LIMIT 10;

-- Reset to the default values
RESET hg_vector_sindi_n_candidate;
RESET hg_vector_sindi_query_prune_ratio;

Recommended configurations by scenario

The following recommendations are measured on the sparse-full dataset, which contains 8.84 million vectors with 30,109 dimensions and 99.6% sparsity.

Scenario

doc_prune_ratio (build)

query_prune_ratio (query)

n_candidate (query)

Expected recall rate

Expected P99 latency

High precision (RAG text retrieval)

0.0–0.2

0.7

500–1000

Greater than 99.8%

About 15 ms

Balanced

0.4

0.8

300–500

Greater than 99.4%

About 13 ms

High performance (real-time search)

0.4–0.6

0.9

500–1000

Greater than 96%

About 10 ms

FAQ

  • Q: Creating a sparse vector index fails with For Sindi algorithm, only 'InnerProduct' distance_method is supported.

    A: The SINDI algorithm supports only the InnerProduct distance method. Change distance_method to InnerProduct. To work with cosine distance, normalize your sparse vectors first and then use InnerProduct.

  • Q: The error If property_key is 'proxima_vectors', only Graph algorithm is supported occurs.

    A: Sparse vector indexes do not support the proxima_vectors property. Use the vectors property to create the index.

  • Q: Writing a sparse vector fails with an error that an index position is out of range.

    A: Index positions in the sparse vector text format (1-based) must be smaller than the term_id_limit value. Check the largest index position in your data and set term_id_limit high enough when you create the index. For example, if the largest index position in your data is 500000, set term_id_limit to a value greater than 500000.

  • Q: Do sparse vector indexes support extra_columns?

    A: No. The SINDI algorithm does not currently support the extra_columns parameter. Do not include it in builder_params.

  • Q: How do I tell whether a query uses the sparse vector index?

    A: Run EXPLAIN ANALYZE to inspect the execution plan. If the plan contains the Vector Filter keyword, the query uses the sparse vector index. If you use the inner_product_distance function for exact search, the plan shows Seq Scan, which means the index is not used.