论文

RAGStress 基准:在知识库被污染的情况下测试 RAG 系统的鲁棒性

RAGStress: A controlled benchmark for evaluating retrieval-augmented generation under knowledge-base degradation

精选理由

做 RAG 的朋友可以看看这个基准,它专门测知识库被污染后系统还能不能扛住,还发现无检索准确率预测不了鲁棒性。

RAGStress 是一个针对 RAG 系统的受控压力测试基准,专门评估知识库受损时的表现。它设计了 4 种污染类型(事实篡改、数字错字、相关性投毒、矛盾注入)和 3 个严重度级别,实验基于 57 个 MMLU 学科的 182,546 篇文档。研究共进行 52,500 次模型-问题-条件评估,发现干净检索会掩盖鲁棒性差异,语义保真类污染比信号效用类扰动危害更大,且无检索时的准确率无法预测受污染检索下的鲁棒性。基准附带生成脚本、污染提示词、元数据模式和评估代码,定位是压力测试工具而非通用排行榜。

原文 · arXiv cs.AI

RAGStress: A controlled benchmark for evaluating retrieval-augmented generation under knowledge-base degradation

Retrieval-Augmented Generation (RAG) is typically evaluated under the implicit assumption that the underlying knowledge base (KB) is clean, leaving the behaviour of RAG systems under realistic KB degradation poorly characterised. We introduce RAGStress, a controlled evaluation benchmark for stress-testing RAG systems under systematic KB corruption. The benchmark pairs four naturalistic corruption types (factual corruption, numeric typo, relevance poisoning, and contradiction injection) with three severity levels (subtle, moderate, and obvious) over a single-KB, metadata-filtered experimental design built from 57 MMLU subjects and 182,546 documents. Across 52,500 model-question-condition evaluations, RAGStress reveals that clean retrieval can mask robustness differences, semantic-fidelity corruptions are substantially more harmful than signal-utility perturbations, no-retrieval accuracy does not predict corrupted-retrieval robustness, and mixed-KB accuracy should not be treated as worst-case robustness. We document the benchmark's intended use, supported claims, and limitations, and provide an artifact bundle including generation scripts, corruption prompts, metadata schema, and evaluation code. RAGStress is intended as a controlled stress test for RAG robustness under KB corruption, not as a general model leaderboard.