Vector analysis
AnalyticDB for PostgreSQLprovides vector analysis capabilities centered on the Nova vector index. It can convert unstructured data such as images, audio, video, and text into feature vectors, and perform vector storage, indexing, and similarity search through SQL. Vector search can also be combined with structured, semi-structured, and full-text search.
Introduction to vector databases
AnalyticDB for PostgreSQL represents unstructured data such as text, images, audio, and video as feature vectors. Based on the MPP architecture and the self-developed next-generation vector search engine Nova, it provides vector storage, indexing, and similarity search capabilities. Nova is the primary vector index, designed for production scenarios such as low-latency writes, multi-tenancy, cross-table hybrid retrieval, and concurrent index building. Users can combine vector search with structured data filtering, semi-structured data queries, and full-text search through SQL, while benefiting from database capabilities such as transactions, high availability, and horizontal scaling. For related core technologies, see the paper "Nova: A Multi-Purpose Vector Engine for Low-Latency, Multi-Tenant, and Cross-Table Hybrid Retrieval".
Vector index selection
AnalyticDB for PostgreSQL provides two types of vector index.
-
Novad (Nova index, disk version): suitable for large-scale data and storage cost-sensitive scenarios.
-
Novam (Nova index, memory version): suitable for low-latency and query performance-sensitive scenarios.
Typical use cases
With AnalyticDB for PostgreSQLvector analysis, you can easily build various intelligent applications.
-
Image search services: retrieve images by using other images as queries.
-
Video retrieval services: retrieve videos by matching specific frame images within videos.
-
Voiceprint retrieval services: match audio clips against other audio clips.
-
Recommendation systems: deliver personalized recommendations by matching user features.
-
Semantic-based text retrieval and recommendation: find semantically similar texts.
-
Q&A bots: build efficient question-answering services by integrating with large language models.
-
File deduplication: remove duplicate files based on file fingerprint features.
Benefits
AnalyticDB for PostgreSQL uses the self-developed vector search engine Nova as its core, deeply integrating vector search with database capabilities. Compared with traditional HNSW indexes, it provides the following key benefits:
-
Hybrid retrieval: supports combining vector search with structured data filtering, semi-structured data queries, and full-text search through SQL to achieve dual-path recall based on semantics and keywords.
-
Adaptive search: automatically balances query speed and recall rate, reducing the need for manual parameter tuning while achieving good retrieval results.
-
Real-time writes: uses a Delta-Base architecture that decouples data writes from primary index building, enabling newly added and updated vectors to participate in search immediately. Automatic backpressure is triggered when data accumulates, and operational capabilities such as Flush are provided.
-
Flexible performance and cost options: provides Novad disk-based indexes for large-scale, cost-sensitive scenarios, and Novam memory-based indexes for low-latency, high-performance scenarios. Supports RaBitQ and SQ8 quantization, as well as various vector types including halfvec, bit, sparsevec, and sq8vector.
-
Online index creation: supports
CREATE INDEX CONCURRENTLY, which allows you to create or rebuild indexes during normal read and write operations, reducing the impact on online business. -
Reduced table lookups: supports using
INCLUDEto add business columns to the vector index. When queries match these columns, results can be returned directly from the index. -
Automatic dimensionality reduction: supports PCA-based automatic dimensionality reduction and manual target dimension specification, reducing the computational cost and storage usage of high-dimensional vectors.
-
Automatic tuning: Autotune supports tuning with real query sets or system-generated synthetic queries. You only need to specify Top-K and the target recall rate, and the system evaluates and applies appropriate search parameters.
-
Easy to use and compatible: you can create tables, import data, create indexes, and perform vector queries through SQL. It also provides a pgvector compatibility mode to reduce migration costs for existing applications.