技巧精选

Milvus HNSW 索引量化解析:SQ、PQ 与 PRQ 如何压缩向量

精选理由

Milvus 官方把 HNSW 的三种量化方案讲清楚了:SQ8/SQ6、PQ、PRQ 各自怎么压缩、省内存多少、代价是什么,做向量检索选型前很值得看。

Milvus 在 HNSW 索引中提供 HNSW_SQ、HNSW_PQ、HNSW_PRQ 三种向量压缩方案。SQ 按维度量化,SQ8 和 SQ6 分别用 8 位和 6 位表示每个维度,SQ4U 采用 4 位无符号值加统一范围。PQ 把向量切成子向量,每个子向量用码本索引表示;PRQ 在 PQ 基础上分阶段编码残差误差,信息保留更多但训练、构建和距离计算成本更高。三者可配合更高精度的 refinement 重排候选以提升召回,建议先定召回目标再对比加载后内存、吞吐和构建时间。

原文 · Milvus

SQ, PQ, or PRQ: how do they compress vectors in Milvus HNSW? HNSW can deliver high recall, but its full precision vectors and layered graph consume substantial memory. Milvus offers HNSW_SQ, HNSW_PQ, and HNSW_PRQ to compress the vectors used in graph construction and distance calculations. Each method encodes vectors differently: • SQ quantizes each dimension. SQ8 and SQ6 use 8 and 6 bits per dimension, with independently estimated ranges. SQ4U uses 4-bit unsigned values and one shared uniform range. • PQ splits the vector into subvectors. Each subvector is represented by a codebook index. • PRQ starts with PQ, then encodes residual error in stages. It retains additional information beyond the initial approximation, with additional training, build, and distance calculation costs. All three quantized index variants use the same HNSW framework, but quantization can change which vectors are connected during construction and how candidates are ranked during retrieval. They can also use higher precision refinement to rerank candidates and potentially improve recall, while adding compute, storage, and memory costs. To compare SQ, PQ, and PRQ, set a recall target first, then measure memory after loading, throughput, and build time—including the cost of any refinement needed to reach that target. 💬 0 🔄 0 ❤️ 0 👀 96 ⚡