Alibaba Cloud Elasticsearch AISearch performance whitepaper

Updated at:

native_hnsw AISearch provides vector search in Alibaba Cloud Elasticsearch. This whitepaper describes the performance test environment, the VectorDBBench test process, and the test results. The tests use Cohere10M as the primary dataset and provide the complete reproduction procedure for a dual 16 vCPU deployment with one primary shard and one replica, as well as reference results for other datasets.

Test environment

Client ECS specifications

Instance typeCPUMemoryNetwork
ecs.g9i.24xlarge96 vCPU384 GiBIn the same VPC as the Elasticsearch instance, using an internal network connection

Server-side Elasticsearch specifications

Product data node specificationsActual ECS specificationsNumber of nodesIndex topology
elasticsearch.turbo2.ga.4xlargeecs.g9i.4xlarge21 primary shard and 1 replica

Before you run the test, disable access logs to prevent logging overhead from affecting the results:

PUT /_cluster/settings
{
  "persistent": {
    "apack.accesslog.enabled": false,
    "apack.accesslog.search.enabled": false,
    "logger.index.search.accesslog": "OFF"
  }
}

Prepare the test tools

The performance evaluation uses the Aliyun Elasticsearch adapter code in zilliztech/VectorDBBench PR #835.

  • Prepare an environment with Python 3.11 or later.

  • Clone VectorDBBench, check out PR #835, and install the Aliyun Elasticsearch dependencies.

git clone https://github.com/zilliztech/VectorDBBench.git
cd VectorDBBench
git fetch origin pull/835/head:pr-835
git checkout pr-835

python3.11 -m venv venv
source venv/bin/activate
python -m pip install -U pip
python -m pip install -e '.[aliyun_elasticsearch]'

Performance test

Test datasets

The tests use Cohere10M as the primary dataset:

DatasetBase vector countVector dimensionsDistance metricNumber of results returned
Cohere10M10,000,000768Cosine SimilarityTop10 / Top100

Set up the dual 16 vCPU test cluster

The scaling test runs on two 16 vCPU data nodes. Each data node holds a complete copy of the Cohere10M graph. Create the instance and configure the ES YML settings in the console before you run the stress test.

Create the instance in the console

  • In the Alibaba Cloud Elasticsearch console, create an instance in Hangzhou Zone J: standard turbo2, 16 vCPU / 64 GiB, two data nodes and one Kibana node, and enable the FalconSeek high-performance search engine.

  • Wait until the instance and both data nodes are in the normal state, and then go to the next section to configure the ES YML settings.

Modify the YML configuration and restart the cluster

Go to the cluster configuration page of the target instance, open the YML configuration editor, and then complete the following steps:

  • In the additional configuration editor at the bottom of the dialog box, paste the following YAML. The platform delivers the configuration to both data nodes at the same time, so you do not need to configure each node separately.

  • When you submit the change, select the in-place change method and confirm the cluster restart.

  • Wait until the change is complete, both data nodes recover, and the target index is green before you start the stress test.

index.native.native_hnsw.search_worker_thread_count: "12"
index.native.native_hnsw.native_search_pool.enabled: "true"
index.native.native_hnsw.native_search_queue_capacity: "208"
index.native.native_hnsw.cpu_lending.mode: "static"

native.indices.queries.cache.enabled: "false"
native.indices.queries.cache.knn.enabled: "false"
indices.queries.cache.size: "0"
Important

Contact us by submitting a ticket to apply the preceding configuration.

This configuration set includes startup settings, so PUT /_cluster/settings cannot replace it. You must apply it through the YML configuration and restart the nodes.

Create the native_hnsw index

Run the following request to create the cohere10m_native_hnsw test index:

PUT /cohere10m_native_hnsw
{
  "settings": {
    "index.havenask.engine.enabled": true,
    "index.sort.field": "id",
    "index.native.index.queries.cache.enabled": false,
    "index.native.index.queries.cache.knn.enabled": false,
    "index.requests.cache.enable": false
  },
  "mappings": {
    "_source": { "enabled": false },
    "properties": {
      "id": { "type": "long" },
      "vector": {
        "type": "dense_vector",
        "dims": 768,
        "index_options": {
          "type": "native_hnsw",
          "builder": {
            "add_chunk_size": 100000,
            "prune_headroom": 0.0
          }
        }
      }
    }
  }
}

The test index uses these native_hnsw graph parameters and contains one primary shard, one replica, and 10,000,000 documents. The primary shard and the replica reside on the two data nodes, and each shard has only one segment.

Query parameters

The Cohere10M tests use the following query parameters:

TopKnumCandidatesrescore oversampleTest concurrency
10904.6212 / 216
1004084.0260 / 64

Run the Cohere10M test from the command line

The following commands run the two-level concurrency throughput test for Top10. Set the connection parameters, username, and password based on the actual configuration of your instance.

set -euo pipefail

export ES_SCHEME='http'
export ES_HOST='<Elasticsearch endpoint>'
export ES_PORT='9200'
export ES_INDEX='cohere10m_native_hnsw'
export ES_USER='<Elasticsearch username>'
export VDBBENCH_ES_QUERY_WIRE_FORMAT='cbor-f32le'
read -rsp 'Elasticsearch password: ' ES_PASSWORD
echo

VDB_COMMON=(
  --scheme "$ES_SCHEME" --host "$ES_HOST" --port "$ES_PORT"
  --user "$ES_USER" --password "${ES_PASSWORD:?ES_PASSWORD is required}"
  --index-name "$ES_INDEX"
  --case-type Performance768D10M
  --skip-drop-old --skip-load
  --m 32 --ef-construction 400
  --skip-search-serial --search-concurrent
  --concurrency-duration 220
)

# Top10: run two concurrency levels, c212 and c216.
vectordbbench aliyunelasticsearch "${VDB_COMMON[@]}" \
  --k 10 --num-candidates 90 \
  --use-rescore --oversample-ratio 4.6 \
  --num-concurrency 212,216 \
  --db-label cohere10m-top10-dual16-c212-c216
Important

The commands include --skip-drop-old and --skip-load, so VectorDBBench skips data loading and searches the existing cohere10m_native_hnsw index. Before you run the test, make sure that the index already contains the 10,000,000 Cohere10M documents. Commands for data loading, the Top100 test, and the serial query runs are not included in this whitepaper.

Test results

Test metrics

The QPS values in the results table use the higher concurrency level and are averaged over two independent 220-second VectorDBBench runs. The Serial P95 and Serial P99 values are obtained by running serial queries separately.

MetricDescription
QPSThe average number of queries per second over a full VectorDBBench run.
Recall@KThe proportion of correct TopK nearest neighbor vectors in the query results.
NDCG@KThe normalized discounted cumulative gain that weights the order of the TopK results.
Serial P9595% of warm serial query latencies are below this value, in milliseconds.
Serial P9999% of warm serial query latencies are below this value, in milliseconds.

native_hnsw search performance

The following table lists the native_hnsw search performance results. The reproduction procedure in this whitepaper covers the Cohere10M dataset. The results for the other datasets are reference results.

DatasetTopKM / efConstructionnumCandidatesRecall / NDCGQPSSerial P95/P99
Cohere10M (768D)1032 / 400900.9797 / 0.980482,520.180.8 / 0.9 ms
Cohere10M (768D)10032 / 4004080.9801 / 0.983321,658.531.9 / 2.0 ms
Cohere1M (768D)1032 / 800980.9901 / 0.990269,083.280.8 / 1.0 ms
Cohere1M (768D)10032 / 8003220.9902 / 0.991924,482.571.5 / 1.7 ms
BioASQ1M (1024D)1016 / 4002440.9504 / 0.951341,701.771.0 / 1.0 ms
BioASQ1M (1024D)10024 / 4003200.9499 / 0.954423,464.421.6 / 1.7 ms
BioASQ10M (1024D)1024 / 4001520.9409 / 0.939946,666.701.0 / 1.1 ms
BioASQ10M (1024D)10024 / 4003200.9381 / 0.942323,927.612.5 / 3.2 ms