arXiv 论文用 3000 道八字选择题评估语言模型的规则应用能力
Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi
用八字命理当考题测模型,六个系统都会背规则但一到大运、事业这类实例题准确率掉到三成多,结论挺有意思。
一篇 arXiv 论文构建了 3000 道中文八字(Bazi)选择题基准,涵盖 14 个理论类别和 11 个案例类别。在 2492 道精炼题目上,六个系统的理论准确率均高于案例准确率,差距在 16.60 至 29.56 个百分点之间。Twelve Stages 和 Nayin 类别平均准确率达 89.10% 和 88.62%,但案例中的 Career 和 Family Relations 仅 36.98% 和 38.19%。DeepSeek 原生配置的对比显示,Flash 和 Pro 的理论成绩分别提升 6.53 和 12.20 分,案例成绩变化为 -3.67 和 +1.27 分。论文指出该基准衡量的是与模型生成答案的一致性,并非真实预测效力。
Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi
Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.