评估主权:元数据驱动分类中的多轨框架与弱监督系统审计

Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems

精选理由

这篇论文戳破了弱监督场景下性能指标的泡沫,做元数据分类或标签质量审计的团队会发现,你报告的准确率可能只是标签流程的镜像,值得点开重新审视评估方法。

AI 摘要

本文提出“评估主权”概念,衡量性能指标独立于标签权威和监督机制的程度。在元数据驱动的弱监督系统中,标签常不完整或不一致,导致模型性能被高估。通过大规模科学元数据的层次多标签分类实验,发现模型在操作环境(银标)下表现良好,但在独立(金标)评估下大幅下降,如Micro-F1从0.54降至0.03。排名指标仍高于基线,表明模型信号与分类有效性存在分歧。研究重新定义评估有效性为系统级属性,并提供审计弱监督系统的实用方法。

原文 · arXiv cs.AI

Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems

Evaluation in machine learning is typically treated as a neutral measurement process. However, in operational information systems, evaluation outcomes are often conditioned by the processes used to generate labels. This paper does not seek to improve classification performance. Instead, it examines the validity of performance measurement under differing label-authority regimes. This issue is particularly relevant in large-scale metadata-driven systems, where labels are often incomplete, inconsistent, or weakly supervised. We introduce evaluation sovereignty, defined as the degree to which performance metrics are independent of label authority and supervision regime, and propose a multi-track evaluation framework that systematically varies training and evaluation label sources. Using hierarchical multi-label classification on large-scale scientific metadata, we demonstrate that models exhibiting strong performance under operational ("silver") evaluation degrade substantially under independent ("gold") evaluation, particularly for fine-grained classification. For example, Micro-F1 decreases from approximately 0.54 to 0.03. Notably, ranking-based metrics remain above baseline, revealing a divergence between latent model signal and classification validity. These findings suggest that commonly reported performance metrics may reflect alignment with labeling processes rather than true predictive capability. We therefore reconceptualize evaluation validity as a system-level property shaped by label governance and provide a practical methodology for auditing intelligent systems operating under weak supervision.