SBW:抗改写的 LLM 智能体行为水印方法
Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents
给智能体加水印,工具改名或观察被改写后老方法基本失效,SBW 把检测率拉到 0.9 以上,还坦白了没防住的攻击。
SBW(Semantic Behavioral Watermarking)针对既有智能体水印的两类弱点提出改进:在语义动作簇层面嵌入水印,并引入密钥化抗碰撞分桶机制,使分桶结果在随机预言机模型下可证明不可预测。在 ToolBench 基准上(每模型 600 条轨迹),改写后的检测率为 0.49-0.66,而精确符号绑定的方法只有 0.05-0.17;在 ALFWorld 上(每模型 100 集)为 0.92-0.97 对 0.00-0.01。密钥化分桶把自适应伪造成功率从 100% 压到误报底线,测试覆盖五个模型(3B-14B,四家厂商)。作者同时公开未解决的边界:链式重放攻击仍能达到 0.76-0.98 的验证率,且抗改写会损失约一半的单步水印容量。
Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents
Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark's own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanović et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B-14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49-0.66 for cluster-level versus 0.05-0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72-83% choice agreement against 22-27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92-0.97 versus 0.00-0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim's own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76-0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at https://anonymous.4open.science/r/SBW-Agent-Watermark.