论文

QCATS:查询上下文感知的 Transformer 切片框架加速数据库内预测查询

QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing

精选理由

一篇数据库方向的论文:把 Transformer 剪成 FFN 切片塞进数据库里跑预测查询,用 BERT-base 和 Qwen-0.6B 实测延迟降到 1/4.42,精度不掉。做数据库+ML 的人可以看看。

arXiv 论文提出 QCATS,一个查询上下文感知的 Transformer 切片框架,用于数据库内的预测查询处理。该方法以查询为粒度,利用查询谓词和元数据统计在模型执行前预选上下文对齐的 FFN 切片,包含离线专家构建和轻量级查询级路由两部分。框架还引入异步 CPU-GPU 流水线和路由感知批处理等系统优化。在四组预测查询工作负载上,QCATS 使用 BERT-base 和 Qwen-0.6B 测试,延迟最多降低 4.42 倍,预测精度与稠密基线相当。

原文 · arXiv cs.AI

QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing

In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.