论文

实证研究:LLM 在日志异常检测中的性能、效率与鲁棒性

Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness

精选理由

一篇把 LLM 做日志异常检测的选型问题拆开讲的论文:哪类适配策略有效、量化掉不掉点、噪声下稳不稳,做运维或 AIOps 的可以看看。

该研究在 3 个公开日志数据集上系统评估了 LLM 用于日志异常检测的表现,覆盖不同适配策略、模型架构、参数规模和量化设置。结果显示不同适配策略带来的性能差异显著,模型扩大的收益因数据集而不同。检测精度相近的模型计算成本差异明显,低比特量化在评估配置下基本保持了检测性能。研究还测试了结构、语义和标签三类噪声在不同扰动强度下的鲁棒性表现。

原文 · arXiv cs.LG

Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness

Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and quantization settings. Our results reveal substantial performance differences across adaptation strategies, while model scaling yields varying detection gains across datasets. We further observe that models with comparable detection accuracy can exhibit markedly different computational costs, and that low-bit quantization largely preserves detection performance in the evaluated configurations. Finally, we examine detection robustness under structural, semantic, and label noise at different perturbation levels. These findings provide empirical insights into the performance, efficiency, and robustness of LLM-based log anomaly detection, highlighting practical considerations beyond conventional accuracy-oriented evaluation.