论文

EADC 基准发布:基于合规知识图谱评测大模型法律合规能力

EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models

精选理由

一份新出的合规评测基准 EADC,用知识图谱加法律专家标注做出 4400 多条 QA,专测模型会不会踩法律红线,做合规的对口。

EADC 是一个面向大模型法律合规评测的新基准,由 AI 合规知识图谱与合规法律专家共同构建。它把抽象法律规则映射为结构化多关系逻辑图,自动生成对抗性合规场景。数据集包含 4,435+ 个问答对,覆盖偏见与歧视、公平性、个人隐私保护、价值观等监管领域。与静态基准不同,它引入上下文长程交互和逻辑驱动风险链,可捕捉绕过传统过滤器的深层合规问题,并揭示了现有先进大模型的监管盲点。

原文 · arXiv cs.AI

EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models

Large Language Models (LLMs) have been used in various industries. However, ensuring their compliance with complex laws and regulatory frameworks remains a great challenge. Existing evaluation paradigms mainly rely on static benchmarks that suffer from three severe limitations: First, the compliance rules being used do not comply with the requirements of Artificial Intelligence (AI) laws and regulations; Second, they only handle apparent, explicit compliance risks, leaving implicit and covert compliance risks undetected; Third, they fail to track the systematic propagation of risks along logical dependency chains or evaluate compliance within nuanced, context-based real-world scenarios. To bridge this critical gap, we introduce EADC, a novel advanced evaluation benchmark of LLMs based on an AI compliance knowledge graph and AI compliance legal experts. By mapping abstract legal rules into structured logical multi-relational graphs, our framework enables automated, evolving agents to distill and synthesize highly sophisticated adversarial scenarios. This compliance benchmark is reviewed and corrected by human AI legal experts throughout the whole process. The resulting dataset (4,435+ QA pairs) provides an extensive, multi-dimensional taxonomy covering critical regulatory frontiers, including bias and discrimination, fairness, personal privacy protection, and values. Crucially, our compliance dataset moves beyond shallow string-matching by incorporating contextual long-horizon interactions and logic-driven hazard chains, capturing deeply embedded compliance anomalies that bypass traditional filters. Experiment evaluations demonstrate that our framework exposes critical regulatory blind spots in state-of-the-art LLMs, offering a rigorous, AI laws and regulations-aligned benchmark to safeguard high-level and deep compliance in the application of LLMs.