Skip to content

feat(milvus): add native BF16 HNSW support - #883

Open
abondarev84 wants to merge 1 commit into
zilliztech:mainfrom
abondarev84:feat/milvus-native-bf16
Open

abondarev84 wants to merge 1 commit into
zilliztech:mainfrom
abondarev84:feat/milvus-native-bf16

Conversation

@abondarev84

Copy link
Copy Markdown

Summary

VectorDBBench creates Milvus dense-vector collections with FLOAT_VECTOR fields, so the benchmark cannot exercise Milvus's native BFLOAT16_VECTOR storage and search path. BF16 is currently reachable only as a scalar-quantizer or refine option (SQType.BF16) over FP32 storage, where the stored field is still 4 bytes per dimension.

This adds an opt-in milvushnswbf16 command that creates a BFLOAT16_VECTOR field and encodes insert and query vectors as BF16 before they reach pymilvus. milvushnsw and all other commands are unchanged.

Design

IndexType.HNSW_BF16 and HNSWBF16Config inherit the existing HNSW options (M, efConstruction, ef, metric handling). The config sends index_type: "HNSW". BF16 is not a separate index algorithm — Milvus selects the BF16 implementation from the field's data type.

_float32_to_bf16_bytes() applies round-to-nearest-even to the IEEE-754 bit patterns and returns one little-endian byte string per vector. NaN inputs remain NaN. It uses NumPy bit operations, so no new runtime dependency is added.

Insert conversion runs once per runner batch. Query conversion runs immediately before the search request, across the whole query batch.

pymilvus accepts raw two-byte BF16 payloads for both insert rows and search placeholders. Verified against 2.6.15, the existing minimum in pyproject.toml. The pymilvus>=2.6.15,<3.0.0 constraint is unchanged.

Request path:

VectorDBBench float32 vectors
  -> NumPy round-to-nearest-even BF16 byte encoding
    -> pymilvus BFLOAT16_VECTOR insert/search request
      -> Milvus BFLOAT16_VECTOR field
        -> Knowhere HNSW BF16 index path

Usage

vectordbbench milvushnswbf16 \
  --case-type Performance768D1M \
  --uri http://localhost:19530 \
  --collection-name bench_bf16 \
  --m 16 \
  --ef-construction 100 \
  --ef-search 64 \
  --k 10

Results

Cohere 1M x 768 (Performance768D1M), M=16 efConstruction=100 ef=64 k=10, COSINE, 1,000,000 vectors per run, 3 runs per configuration. Both hosts: 8 vCPU / 16 GB, Milvus 3.0.0. One instance per architecture for all runs. Unpatched main and this branch, from separate checkouts.

recall@10: fraction of the dataset's 10 true nearest neighbors present in the 10 returned results, averaged over 1,000 queries. Range 0 to 1. Higher is better.

Host Command Field type recall@10, mean of 3 (min–max)
x86-64 milvushnsw (unpatched main) FLOAT_VECTOR 0.9317 (0.9298–0.9328)
x86-64 milvushnswbf16 (this branch) BFLOAT16_VECTOR 0.9305 (0.9295–0.9315)
aarch64 milvushnsw (unpatched main) FLOAT_VECTOR 0.9294 (0.9284–0.9299)
aarch64 milvushnswbf16 (this branch) BFLOAT16_VECTOR 0.9306 (0.9294–0.9315)

Add IndexType.HNSW_BF16, HNSWBF16Config, and a MilvusHNSWBF16 CLI
command. The client creates BFLOAT16_VECTOR fields and converts
float32 insert batches and query vectors to round-to-nearest-even
BF16 byte payloads accepted by pymilvus.

The conversion uses NumPy bit operations, avoiding an additional
ml_dtypes runtime dependency. Milvus still receives index_type
"HNSW"; BF16 is selected by the collection field's data type, so
FP32, FTS, and GPU configurations keep their existing behavior.

Tests cover the rounding and special-value encoding, BFLOAT16_VECTOR
schema selection against the FP32 default, insert and batch-search
conversion, the CLI command wiring, and HNSW_BF16 case-config
registration.

Co-authored-by: RJ Silk <robesilk@amazon.com>
@sre-ci-robot

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: abondarev84
To complete the pull request process, please assign xuanyang-cn after the PR has been reviewed.
You can assign the PR to them by writing /assign @xuanyang-cn in a comment when ready.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@abondarev84

Copy link
Copy Markdown
Author

/assign @XuanYang-cn

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants