论文多源确认

JEV 与 GPT-6 Luna、Qwen3.8-27B 在政治学文本标注任务上的对比

JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications

精选理由

有人测了 TypeSafe 那个号称便宜快的 JEV,结果在 OpenAI 批量价格下根本不省钱,校准也没明显赢 Qwen3.8-27B,想做文本标注的可以看看。

论文在七项政治学复制研究上测试了 TypeSafe 推出的 JEV 模型与 LLM 的表现。结果显示 JEV 在多项任务上接近 GPT-6 Luna 和开源的 Qwen3.8-27B,但在 OpenAI 批量价格下没有成本优势。单次提问时 JEV 的概率校准优于 GPT-6 Luna,但不稳定地好于 Qwen3.8-27B。作者的结论是 JEV 的主要优势在于速度快和解析选择概率方便。

原文 · arXiv: OpenAI

JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications

Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.