自适应分类器的盲区:内部更新如何让漂移检测器失效
The Blind Spot Paradox: When Adaptive Classifiers Defeat Drift Detectors
一篇挺反直觉的论文:模型自己修复得太快,漂移检测器反而看不见漂移了,还给出 Δe=0.120 的检测下限,做流式监控的值得看实验设计。
论文指出用自适应分类器的错误流做概念漂移检测存在结构性冲突:内部适应快于证据累积,CUSUM、Page-Hinkley 等累积检测器到不了阈值。对 Adaptive Random Forest(ARF)的实测显示,存活树通过增量叶子更新吸收了 98.6% 的漂移后错误瞬态,而第一次后台树替换仅占 0.71% 的被抹除错误量,却让外部检测率下降 31 个百分点。作者推导出累积证据无法过阈值的有限时域边界,并测得临界幅度下限 Δe_c=0.120。验证覆盖合成漂移、ARMA-GARCH 序列(ProteuS)和 BAF、INSECTS 表格基准;在标准阈值下的合成扫描中,盲区出现在 Δe≈0.25,这也解释了 SEA(Δe≤0.21)等经典基准为何测不到它。
The Blind Spot Paradox: When Adaptive Classifiers Defeat Drift Detectors
Monitoring concept drift from an adaptive classifier's error stream creates an operational conflict with the model's own update loop. When internal adaptation outpaces evidence accumulation, accuracy recovers before cumulative detectors (CUSUM, Page-Hinkley) can reach threshold. Instrumenting an Adaptive Random Forest (ARF) shows that surviving trees absorb 98.6% of the post-drift error transient through incremental leaf updates alone. The first background tree swap accounts for just 0.71% of this erased error volume, but drops external detection rates by 31 percentage points. We derive the finite-horizon boundary where cumulative evidence fails to cross threshold and measure a critical magnitude floor ($Δe_c = 0.120$) below which false-alarm budgets preclude detection. This failure manifests as missed shifts on stationary streams and false-alarm flooding triggered by internal tree swaps on noisy baselines. We validate on synthetic shifts, ARMA-GARCH series (ProteuS), and tabular benchmarks (BAF, INSECTS); on the synthetic sweep at a standard threshold, the blind spot appears at $Δe \approx 0.25$, showing why classical benchmarks like SEA ($Δe \le 0.21$) failed to reach it.