论文精选

K-Bench:为高风险心理健康对话开发临床校准基准

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

精选理由

朋友推荐:新开发的K-Bench基准专门用来测试AI在处理高风险心理健康对话时的表现,能帮你了解不同模型在应对自杀、自残等场景时的能力差异。

研究人员开发了K-Bench基准,用于评估大语言模型在涉及自杀、自残等高风险心理健康对话中的表现。该基准包含200个多轮场景,涵盖14家供应商的33个基础模型配置。一个冻结的GPT-4o判别器与临床专家共识的准确率高达94.2%。领先模型在支持性对话和风险评分方面表现优异,而风险探索则暴露出低性能配置的显著差异。

原文 · arXiv cs.AI

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

% !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading models combined strong supportive conversation with combined-risk scores above 95, whereas risk exploration exposed substantial variation among lower-performing configurations. Therapeutic prompting produced configuration-specific gains concentrated among weaker models, while elevated reasoning produced no average improvement. K-Bench combines broader clinical coverage and configuration-scale comparison with a continuously updated public leaderboard whose operational test materials are protected from direct optimisation. The leaderboard is available at www.k-bench.ai.