Jev 科学决策评测:12 个配置对比语义选择,延迟最低
Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
有人把 Jev 当科学决策组件实测了一遍:20 道语义题跟 12 个配置比,它全对且延迟最低。
arXiv 论文将 Jev 作为科学工作流中的语义决策组件进行评估,测试框架按其文档指引运行,并把算术运算交由代码执行。实验在十个科学案例的二十个 Choices 上对比十二个模型配置,每个任务重复五次,分别统计语义选择、下游输出和最终结论标签。结果显示 Jev 与另外五个配置在语义选择上达到完全正确,并在成功响应中取得最低的中位延迟。研究还发现,三个对比模型在一个文化历史问题上出现七次错误选择,导致下游计数变化但最终结论标签仍保持正确。
Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning of the resulting count or comparison. We evaluate Jev as a semantic decision component using a harness that follows its documented guidance and assigns arithmetic to code. The study compares twelve model configurations on twenty source-grounded Choices across ten scientific cases, each repeated five times. We measure semantic selections, downstream outputs and final claim labels separately. Jev matched five other configurations at complete semantic correctness and achieved the lowest observed median latency among successful responses. Across three comparison models, seven wrong selections on one culture-history question changed downstream counts while preserving the correct final label. These results identify a useful role for Jev in prepared scientific decision tasks and show why evaluating that role requires checking the relations and quantities that a workflow will reuse.