QUILT优化稀疏注意力预填充
QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
QUILT通过共享查询执行优化稀疏注意力,在GLM-5.3和DeepSeek-3.2上实现55%延迟降低,比现有方案更高效。
QUILT是一种新型稀疏注意力执行机制,通过联合处理邻近查询和重用共享KV条目,减少冗余内存流量和计算。该研究引入Shift-and-Compare Set Decomposition(SCSD)技术,将不规则集合操作转换为适合现代加速器的规则数据并行原语。在LongBench基准测试中,使用GLM-5.3和DeepSeek-3.2模型,QUILT将平均内核延迟降低最多55.1%,处理KV数据减少最多55.9%,同时将首次输出时间(TTFT)延迟降低最多36.8%,且精度损失可忽略不计。
QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries. We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse. We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation. QUILT introduces Shift-and-Compare Set Decomposition (SCSD), which transforms irregular set operations into regular data-parallel primitives suitable for modern accelerators, and pipelines SCSD with attention computation to hide its overhead. Cascaded sharing captures reuse hierarchically at multiple granularities. A tile-aware execution strategy balances sharing granularity with hardware tile utilization and selectively removes low-importance query-specific tails to eliminate underutilized tiles. We evaluate QUILT on LongBench using GLM-5.3 and DeepSeek-3.2 under both tensor and sequence parallelism. Compared with the state-of-the-art sparse-attention kernel, QUILT reduces average kernel latency by up to 55.1% and processed KV data by up to 55.9%, while reducing time-to-first-token (TTFT) latency by up to 36.8% with negligible accuracy degradation.