BackTrend 基准发布:用反向重构评估科学弱信号预测
BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction
新论文 BackTrend 让模型回猜哪些冷门研究会火,最强系统 F1 只有 10.1%,AI 前瞻能力终于有了硬标尺。
arXiv 论文提出 BackTrend 基准,要求系统在给定成熟目标话题和时间证据约束下,回溯找出未被重视的研究问题与新兴方法两类前兆信号。基准覆盖人工智能与机器学习领域的 25 个成熟目标话题和 66 条经人工验证的弱信号,每条信号锚定其 2019-2024 年的发表频率轨迹。团队用语义匹配和覆盖率指标评估前沿 LLM、RAG 系统与智能体研究系统,最强系统 F1 仅 10.1%,Coverage10 最多覆盖 18.5% 的参考信号。预算分析显示,增加检索和网络搜索证据在中低预算区间能改善表现,但单靠更多证据无法填平这一差距。
BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction
Scientific weak signals are early, low-visibility research directions that later become central to mature scientific topics, yet existing resources such as trend tracking, citation forecasting, and foresight reports rarely provide validated reference sets that link concrete early precursors to later paradigms. We introduce BackTrend, a retrospective benchmark in which, given a mature target topic and a temporal evidence constraint, systems must recover two types of precursors: problem-space signals, underrecognized research problems, and solution-space signals, emerging methods for known problems. BackTrend contains 25 mature target topics in artificial intelligence and machine learning and 66 human-validated weak signals, reconstructed from large-scale literature by grounding each candidate in its 2019-2024 publication-frequency trajectory. We evaluate frontier LLMs, RAG systems, and agentic research systems using semantic matching and coverage-based metrics. Current systems often generate plausible but misaligned precursors, exhibiting topic drift, granularity mismatch, near-miss matching, and incomplete coverage; the strongest system achieves only 10.1% F1, while Coverage10 reaches at most 18.5% of the reference signals. Our budget analyses show that additional retrieval and web-search evidence can improve performance up to a moderate budget, but does not by itself close the substantial performance gap.