论文

协议选择如何影响日志异常检测基准结论

Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

精选理由

做日志异常检测的可以看看,同一套检测器换个评估协议,HDFS 和 BGL 上的排名就能变,选模型前先想清楚测试集怎么切。

研究在 HDFS 与 BGL 日志上用六种固定配置测试切分方式、表征可见性与组件成本对基准结果的影响。随机切分会让多个配置接近平均精度上限,而 HDFS 组不相交切分与 BGL 按时间顺序评估得分更低且排名不同。在固定 BGL 截止点上,解析器选择使语义 XGBoost 的平均精度波动 0.124。HDFS 到 BGL 的跨系统迁移中,加入目标模板的 union-corpus 访问使平均精度从 0.191 升至 0.325。

原文 · arXiv cs.AI

Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, sequence, and semantic configurations on Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) logs. Random splits place several configurations near the average-precision ceiling, whereas group-disjoint HDFS and chronological BGL evaluation produce lower scores and different observed orderings. At a fixed BGL cutoff, parser choice spans 0.124 in semantic XGBoost mean average precision while preserving its lead over count XGBoost; the earliest rolling period reverses that ordering. A two-factor cross-system ablation contrasts source-only representations with offline transductive access to unlabeled target templates through the representation corpus and inverse document frequency: HDFS-to-BGL mean average precision moves from 0.191 with source-only access to 0.325 with union-corpus, target-IDF access, and the intermediate conditions reveal direction-dependent interactions in average precision and retrieval at fixed review budgets. Component-level profiling separates parsing and representation costs from classifier training, prediction, and storage. Together, these findings connect detector comparisons to the test population, preprocessing state, visible information, and measured pipeline stages, and identify the protocol fields needed alongside a score to support interpretable comparisons of log anomaly detection accuracy and resource use.