OSS Vectors quick start

Updated at:

This topic shows how to quickly implement a complete application workflow, from data preparation to vector retrieval. The process involves four core steps: creating a vector bucket, creating a vector index, uploading vector data, and performing vector retrieval.

Before you begin, make sure that:

  • You have activated Object Storage Service (OSS).

  • Standard-mode is currently available in the following regions: China (Shenzhen), China (Qingdao), China (Beijing), China (Hangzhou), China (Shanghai), China (Ulanqab), Singapore, China (Hong Kong), Indonesia (Jakarta), Germany (Frankfurt), US (Silicon Valley), and US (Virginia).

  • Fusion-mode is currently in invitational preview and available only in the Indonesia (Jakarta) region.

Step 1: Create a vector bucket

Create a vector bucket to contain all your vector data and indexes.

  1. On the Vector Buckets page, click Create a Vector Bucket.

  2. Configure the bucket information:

    • Vector Bucket Name: Enter a name for the bucket that is unique to your Alibaba Cloud account within a region. The name must be 3 to 32 characters in length and can contain only lowercase letters, digits, and hyphens (-). It cannot start or end with a hyphen.

    • Region: Select the region for your business, for example, "China (Shenzhen)".

  3. Click OK to create the bucket.

Step 2: Create a vector index

After creating the bucket, create a vector index within it. The index defines the structure of your vectors, such as their dimension, and the retrieval method, such as the distance metric. It is the foundation for storing and querying vector data. Vector indexes come in two modes: Standard-mode and Fusion-mode. You can create an index by using the console, the OSS SDK, ossutil, or the API.

When you configure the index parameters, regardless of the mode, you must first specify an index table name and then select an index mode. Then, configure the key parameters for the mode that you selected:

  • Index Table Name: Enter a name for the index that is unique within the bucket. The name must be 1 to 63 characters in length, consist of letters and digits, and start with a letter.

  • Index Mode: Select Standard-mode or Fusion-mode based on your use case.

The following table compares the key parameters and creation methods of the two modes:

Item

Standard-mode

Fusion-mode

Key parameters

Vector Data Type: The default value is float32 (32-bit floating-point).

Vector Dimension: Set the dimension of the vectors, for example, 128. The value must be an integer from 1 to 4,096. All vectors that you upload to this index must have the same dimension.

Distance Metric Function: Select a distance calculation method based on your use case. Euclidean distance is ideal for measuring numerical differences, and cosine distance is suitable for calculating semantic similarity for high-dimensional data such as text and images.

Field schema: Fusion-mode uses an explicit schema, so you must define the field structure. A single index table supports a maximum of 3 vector fields, and the total number of vector and scalar fields cannot exceed 100. At least one field of the vector type is required.

Vector fields: Specify the vector dimension and the distance metric function. Euclidean distance, cosine distance, and maximum inner product are supported.

Scalar fields: Define fields of types such as string, long, double, bool, ip, and geoPoint as needed for capabilities such as full-text search and scalar filtering.

Creation method

You can create the index directly in the console: on the Vector Buckets page, click the name of the vector bucket that you created, go to the Vector Indexes page, click Create Index Table, configure the key parameters, and then click OK. You can also use the OSS SDK, ossutil, or the API.

This mode cannot be created in the console yet. To create an index, call the PutVectorIndexFusion operation, or use ossutil or the OSS SDK.

Step 3: Upload vector data

Once the index is ready, you can upload vector data to it.

  1. In the index list, find the index that you just created and click View Data on the right.

  2. On the index page, click Vector Data Insertion.

  3. Configure the vector data. You can add multiple vector entries at once:

    • Primary Key Value: Set a unique identifier for the vector.

    • Vector Data: Enter the vector values as a comma-separated list of numbers. The dimension of the vector (the number of values) must exactly match the Vector Dimension set in Step 2.

    • Metadata: You can add optional metadata, such as categories, titles, or timestamps. Metadata enables precise filtering during retrieval.

  4. Click OK to insert the data.

Step 4: Perform vector retrieval

After you prepare the data, you can perform vector retrieval, which is the core part of the workflow. Typically, you call the API from your application by using an SDK to perform vector retrieval and quickly locate the target data.

The following Python SDK example shows how to retrieve the top 10 data entries that are most similar to a target vector and whose type field is not "comedy" or "documentary".

import argparse
import alibabacloud_oss_v2 as oss
import alibabacloud_oss_v2.vectors as oss_vectors

parser = argparse.ArgumentParser(description="vector query vectors sample")
parser.add_argument('--region', help='The region in which the bucket is located.', required=True)
parser.add_argument('--bucket', help='The name of the bucket.', required=True)
parser.add_argument('--endpoint', help='The domain names that other services can use to access OSS')
parser.add_argument('--index_name', help='The name of the vector index.', required=True)
parser.add_argument('--account_id', help='The account id.', required=True)

def main():
    args = parser.parse_args()

    # Loading credentials values from the environment variables
    credentials_provider = oss.credentials.EnvironmentVariableCredentialsProvider()

    # Using the SDK's default configuration
    cfg = oss.config.load_default()
    cfg.credentials_provider = credentials_provider
    cfg.region = args.region
    cfg.account_id = args.account_id
    cfg.use_internal_endpoint = True  # To access the service over the public network, set this to False or remove this line.
    if args.endpoint is not None:
        cfg.endpoint = args.endpoint

    vector_client = oss_vectors.Client(cfg)

    query_filter = {
        "$and": [{
            "type": {
                "$nin": ["comedy", "documentary"]
            }
        }]
    }

    query_vector = {"float32": [0.1] * 128}

    result = vector_client.query_vectors(oss_vectors.models.QueryVectorsRequest(
        bucket=args.bucket,
        index_name=args.index_name,
        filter=query_filter,
        query_vector=query_vector,
        return_distance=True,
        return_metadata=True,
        top_k=10
    ))

    print(f'status code: {result.status_code},'
          f' request id: {result.request_id},'
          )

    if result.vectors:
        for vector in result.vectors:
            print(f'vector: {vector}')


if __name__ == "__main__":
    main()

If you use a Fusion-mode index table, call the QueryVectorsFusion operation to perform hybrid retrieval. The following Python SDK example combines a vector nearest-neighbor search (knn) with a full-text search on the title field, applies a scalar filter on the year field (greater than or equal to 2020), and returns the top 10 most relevant results.

import argparse
import base64
import hashlib
import json

from requests.structures import CaseInsensitiveDict

import alibabacloud_oss_v2 as oss
import alibabacloud_oss_v2.vectors as oss_vectors
from alibabacloud_oss_v2 import OperationInput


parser = argparse.ArgumentParser(description="query vectors fusion sample")
parser.add_argument('--region', help='The region in which the bucket is located.', required=True)
parser.add_argument('--bucket', help='The name of the vector bucket.', required=True)
parser.add_argument('--endpoint', help='The domain names that other services can use to access OSS')
parser.add_argument('--index_name', help='The name of the fusion vector index.', required=True)
parser.add_argument('--account_id', help='The account id.', required=True)


def main():
    args = parser.parse_args()

    # Loading credentials values from the environment variables
    credentials_provider = oss.credentials.EnvironmentVariableCredentialsProvider()

    # Using the SDK's default configuration
    cfg = oss.config.load_default()
    cfg.credentials_provider = credentials_provider
    cfg.region = args.region
    cfg.account_id = args.account_id
    if args.endpoint is not None:
        cfg.endpoint = args.endpoint

    client = oss_vectors.Client(cfg)

    # Query vectors fusion
    json_data = json.dumps({
        'indexName': args.index_name,
        'knn': {
            'field': 'text_vector',
            'queryVector': [0.1, 0.2, 0.3, 0.4],
            'topK': 10,
            'numCandidates': 100,
            'filter': {
                'year': {
                    '$gte': 2020,
                },
            },
        },
        'query': {
            'title': {
                '$textMatch': {
                    'value': 'vector search',
                    'boost': 2.0,
                },
            },
        },
        'returnMetadata': True,
        'returnMetadataFields': ['title', 'category', 'year'],
        'limit': 10,
        'sort': [
            {
                '_score': {
                    'order': 'desc',
                },
            },
        ],
    })
    body = json_data.encode()
    content_md5 = base64.b64encode(hashlib.md5(body).digest()).decode()
    op_input = OperationInput(
        op_name='QueryVectorsFusion',
        method='POST',
        headers=CaseInsensitiveDict({
            'Content-Type': 'application/json',
            'Content-MD5': content_md5,
        }),
        parameters={
            'queryVectorsFusion': '',
        },
        bucket=args.bucket,
        body=body,
    )
    op_output = client.invoke_operation(op_input)
    print(f'QueryVectorsFusion status code: {op_output.status_code},'
          f' request id: {op_output.headers.get("x-oss-request-id", "")},'
          f' content: {op_output.http_response.content},')


if __name__ == "__main__":
    main()

Next steps

You can perform all operations on vector buckets by using the console, the OSS SDK, ossutil, or direct API calls. This quick start shows the fastest way to get started. For details on advanced configurations and usage, see the following topics: