论文

SeLATM 论文提出分段级智能体主题建模框架,降低 LLM 资源消耗

Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency

精选理由

一篇处理长文档主题建模的论文,把文档切段再生成主题,用智能体反馈细化,比逐篇分配主题省不少 token,做文本分析可以看看。

arXiv 论文提出 SeLATM 框架,针对 LLM 主题建模在文档主题分配中的三个问题:无法生成文档级主题分布、主题过宽或过窄、随文档数量和长度增长的高资源消耗。SeLATM 采用分段级主题生成,并通过智能体反馈循环进行主题细化。在多个数据集上的实验显示,相比基于主题分配的方法,SeLATM 显著减少 LLM 资源消耗,同时保持更好的性能。该框架面向需要处理大量文档的工业级文本挖掘场景。

原文 · arXiv cs.AI

Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency

Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as the incapability to produce topic distributions over a document, too broad or narrow topics, and high resource consumption, which increases with the number and length of of documents being assigned topics. These issues are particularly critical for industrial applications, which require high-quality, in-depth analysis and the processing of large volumes of documents. In this context, this paper introduces a framework called SeLATM, which addresses these concerns by employing segment-level topic generation and topic refinement through agentic feedback loops. Experimental results on various datasets demonstrate that SeLATM significantly reduces the LLM resources compared to methods based on topic assignment process, while maintaining superior performance.