生物循环:基于加速自适应的 CRISPR 筛选命中发现
Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens
这个研究挺有意思,他们用 LLM 帮助做 CRISPR 筛选,比随机选好很多,能更快找到关键基因。
本文提出 AssayBench-Loop 大规模基准,包含 1,389 个 CRISPR 筛选实验。基于此,他们开发了 AssayLoop 框架,结合了基于 LLM 的生物先验知识,在 5% 的候选库测试中实现了 5.67 倍的随机选择提升,并成功恢复 27.7% 的命中。
Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens
Many biological discovery problems require experiments to be selected sequentially under constrained budgets. CRISPR screening is a prominent example, as exhaustive perturbation testing is often infeasible and candidate perturbations must instead be prioritized over multiple experimental rounds. Despite the importance of this problem, existing benchmarks for adaptive hit discovery remain limited in scale and diversity. Here, we introduce AssayBench-Loop, a large-scale benchmark for adaptive hit discovery comprising 1,389 CRISPR screens across five phenotype categories. Beyond enabling systematic evaluation, its scale makes it possible to learn acquisition strategies across historical experiments. Building on this resource, we introduce AssayLoop, a sequential experimental design framework combining AssayFormer, a transformer-based amortized acquisition policy trained across historical screens to adapt from experimental feedback, with LLM-derived biological priors through an adaptive handoff. In this view, completed experiments become training data for learning how accumulated evidence should guide what to test next, while LLMs provide prior biological knowledge to seed the search. We further introduce AssayLLM, showing that the same principle can be extended directly to an LLM through task-specific post-training. On temporally held-out screens, AssayLoop achieves a 5.67-fold enrichment over random selection and recovers 27.7% of hits after assaying approximately 5% of the candidate library, outperforming existing adaptive-design methods and standalone LLMs, and AssayFormer alone. Performance improves with increasing historical training data and transfers to phenotype categories excluded from training. These results demonstrate the value of learning acquisition policies across historical experiments and combining them with broad biological priors for efficient adaptive hit discovery.