Alibaba Cloud Elasticsearch AISearch performance whitepaper
native_hnsw AISearch provides vector search in Alibaba Cloud Elasticsearch. This whitepaper describes the performance test environment, the VectorDBBench test process, and the test results. The tests use Cohere10M as the primary dataset and provide the complete reproduction procedure for a dual 16 vCPU deployment with one primary shard and one replica, as well as reference results for other datasets.
Test environment
Client ECS specifications
| Instance type | CPU | Memory | Network |
ecs.g9i.24xlarge | 96 vCPU | 384 GiB | In the same VPC as the Elasticsearch instance, using an internal network connection |
Server-side Elasticsearch specifications
| Product data node specifications | Actual ECS specifications | Number of nodes | Index topology |
elasticsearch.turbo2.ga.4xlarge | ecs.g9i.4xlarge | 2 | 1 primary shard and 1 replica |
Before you run the test, disable access logs to prevent logging overhead from affecting the results:
PUT /_cluster/settings
{
"persistent": {
"apack.accesslog.enabled": false,
"apack.accesslog.search.enabled": false,
"logger.index.search.accesslog": "OFF"
}
}Prepare the test tools
The performance evaluation uses the Aliyun Elasticsearch adapter code in zilliztech/VectorDBBench PR #835.
Prepare an environment with Python 3.11 or later.
Clone VectorDBBench, check out PR #835, and install the Aliyun Elasticsearch dependencies.
git clone https://github.com/zilliztech/VectorDBBench.git
cd VectorDBBench
git fetch origin pull/835/head:pr-835
git checkout pr-835
python3.11 -m venv venv
source venv/bin/activate
python -m pip install -U pip
python -m pip install -e '.[aliyun_elasticsearch]'Performance test
Test datasets
The tests use Cohere10M as the primary dataset:
| Dataset | Base vector count | Vector dimensions | Distance metric | Number of results returned |
| Cohere10M | 10,000,000 | 768 | Cosine Similarity | Top10 / Top100 |
Set up the dual 16 vCPU test cluster
The scaling test runs on two 16 vCPU data nodes. Each data node holds a complete copy of the Cohere10M graph. Create the instance and configure the ES YML settings in the console before you run the stress test.
Create the instance in the console
In the Alibaba Cloud Elasticsearch console, create an instance in Hangzhou Zone J: standard turbo2, 16 vCPU / 64 GiB, two data nodes and one Kibana node, and enable the FalconSeek high-performance search engine.
Wait until the instance and both data nodes are in the normal state, and then go to the next section to configure the ES YML settings.
Modify the YML configuration and restart the cluster
Go to the cluster configuration page of the target instance, open the YML configuration editor, and then complete the following steps:
In the additional configuration editor at the bottom of the dialog box, paste the following YAML. The platform delivers the configuration to both data nodes at the same time, so you do not need to configure each node separately.
When you submit the change, select the in-place change method and confirm the cluster restart.
Wait until the change is complete, both data nodes recover, and the target index is green before you start the stress test.
index.native.native_hnsw.search_worker_thread_count: "12"
index.native.native_hnsw.native_search_pool.enabled: "true"
index.native.native_hnsw.native_search_queue_capacity: "208"
index.native.native_hnsw.cpu_lending.mode: "static"
native.indices.queries.cache.enabled: "false"
native.indices.queries.cache.knn.enabled: "false"
indices.queries.cache.size: "0"Contact us by submitting a ticket to apply the preceding configuration.
This configuration set includes startup settings, so PUT /_cluster/settings cannot replace it. You must apply it through the YML configuration and restart the nodes.
Create the native_hnsw index
Run the following request to create the cohere10m_native_hnsw test index:
PUT /cohere10m_native_hnsw
{
"settings": {
"index.havenask.engine.enabled": true,
"index.sort.field": "id",
"index.native.index.queries.cache.enabled": false,
"index.native.index.queries.cache.knn.enabled": false,
"index.requests.cache.enable": false
},
"mappings": {
"_source": { "enabled": false },
"properties": {
"id": { "type": "long" },
"vector": {
"type": "dense_vector",
"dims": 768,
"index_options": {
"type": "native_hnsw",
"builder": {
"add_chunk_size": 100000,
"prune_headroom": 0.0
}
}
}
}
}
}The test index uses these native_hnsw graph parameters and contains one primary shard, one replica, and 10,000,000 documents. The primary shard and the replica reside on the two data nodes, and each shard has only one segment.
Query parameters
The Cohere10M tests use the following query parameters:
| TopK | numCandidates | rescore oversample | Test concurrency |
| 10 | 90 | 4.6 | 212 / 216 |
| 100 | 408 | 4.02 | 60 / 64 |
Run the Cohere10M test from the command line
The following commands run the two-level concurrency throughput test for Top10. Set the connection parameters, username, and password based on the actual configuration of your instance.
set -euo pipefail
export ES_SCHEME='http'
export ES_HOST='<Elasticsearch endpoint>'
export ES_PORT='9200'
export ES_INDEX='cohere10m_native_hnsw'
export ES_USER='<Elasticsearch username>'
export VDBBENCH_ES_QUERY_WIRE_FORMAT='cbor-f32le'
read -rsp 'Elasticsearch password: ' ES_PASSWORD
echo
VDB_COMMON=(
--scheme "$ES_SCHEME" --host "$ES_HOST" --port "$ES_PORT"
--user "$ES_USER" --password "${ES_PASSWORD:?ES_PASSWORD is required}"
--index-name "$ES_INDEX"
--case-type Performance768D10M
--skip-drop-old --skip-load
--m 32 --ef-construction 400
--skip-search-serial --search-concurrent
--concurrency-duration 220
)
# Top10: run two concurrency levels, c212 and c216.
vectordbbench aliyunelasticsearch "${VDB_COMMON[@]}" \
--k 10 --num-candidates 90 \
--use-rescore --oversample-ratio 4.6 \
--num-concurrency 212,216 \
--db-label cohere10m-top10-dual16-c212-c216The commands include --skip-drop-old and --skip-load, so VectorDBBench skips data loading and searches the existing cohere10m_native_hnsw index. Before you run the test, make sure that the index already contains the 10,000,000 Cohere10M documents. Commands for data loading, the Top100 test, and the serial query runs are not included in this whitepaper.
Test results
Test metrics
The QPS values in the results table use the higher concurrency level and are averaged over two independent 220-second VectorDBBench runs. The Serial P95 and Serial P99 values are obtained by running serial queries separately.
| Metric | Description |
| QPS | The average number of queries per second over a full VectorDBBench run. |
| Recall@K | The proportion of correct TopK nearest neighbor vectors in the query results. |
| NDCG@K | The normalized discounted cumulative gain that weights the order of the TopK results. |
| Serial P95 | 95% of warm serial query latencies are below this value, in milliseconds. |
| Serial P99 | 99% of warm serial query latencies are below this value, in milliseconds. |
native_hnsw search performance
The following table lists the native_hnsw search performance results. The reproduction procedure in this whitepaper covers the Cohere10M dataset. The results for the other datasets are reference results.
| Dataset | TopK | M / efConstruction | numCandidates | Recall / NDCG | QPS | Serial P95/P99 |
| Cohere10M (768D) | 10 | 32 / 400 | 90 | 0.9797 / 0.9804 | 82,520.18 | 0.8 / 0.9 ms |
| Cohere10M (768D) | 100 | 32 / 400 | 408 | 0.9801 / 0.9833 | 21,658.53 | 1.9 / 2.0 ms |
| Cohere1M (768D) | 10 | 32 / 800 | 98 | 0.9901 / 0.9902 | 69,083.28 | 0.8 / 1.0 ms |
| Cohere1M (768D) | 100 | 32 / 800 | 322 | 0.9902 / 0.9919 | 24,482.57 | 1.5 / 1.7 ms |
| BioASQ1M (1024D) | 10 | 16 / 400 | 244 | 0.9504 / 0.9513 | 41,701.77 | 1.0 / 1.0 ms |
| BioASQ1M (1024D) | 100 | 24 / 400 | 320 | 0.9499 / 0.9544 | 23,464.42 | 1.6 / 1.7 ms |
| BioASQ10M (1024D) | 10 | 24 / 400 | 152 | 0.9409 / 0.9399 | 46,666.70 | 1.0 / 1.1 ms |
| BioASQ10M (1024D) | 100 | 24 / 400 | 320 | 0.9381 / 0.9423 | 23,927.61 | 2.5 / 3.2 ms |