使用Gensim进行文本摘要时遇到短段落处理问题
解决Gensim summarize处理短段落的警告与无摘要问题
嘿,我来帮你搞定这个问题!你遇到的警告其实已经把原因说得明明白白——Gensim的summarize函数默认要求输入文本至少包含10个句子,因为它背后依赖的TextRank算法需要足够的句子节点来构建语义关联图,才能选出重要句子生成摘要。你的输入只有1个短句子,完全达不到这个要求,所以才会弹出警告,还生成不了摘要。
下面给你两个实用的解决方案,按需选择:
方案1:给模型提供足够长度的文本
如果你的场景允许使用长文本,只要输入包含至少10个句子,summarize就能正常工作。比如这样修改代码:
import logging logging.basicConfig(format='%(asctime)s : %(levelname)s : %(message)s', level=logging.INFO) from gensim.summarization import summarize # 示例长文本(包含多个句子) text = """Natural language processing (NLP) is a subfield of linguistics, computer science, and artificial intelligence concerned with the interactions between computers and human language. It focuses on how to program computers to process and analyze large amounts of natural language data. Challenges in natural language processing frequently involve speech recognition, natural language understanding, and natural language generation. Gensim is a Python library designed for efficient processing of text data, especially for topic modeling, document indexing, and similarity retrieval. The summarize module in Gensim uses the TextRank algorithm to extract important sentences from a text. TextRank is an unsupervised learning algorithm inspired by PageRank, which is used by Google to rank web pages. It works by constructing a graph where each node represents a sentence, and edges represent the similarity between sentences. The algorithm then assigns a score to each sentence based on the connections it has to other sentences. Sentences with higher scores are considered more important and are included in the summary.""" print('Summary:') print(summarize(text))
方案2:适配短文本的替代方案
如果你的需求就是处理短段落,那Gensim的summarize确实不太合适,试试下面两种思路:
思路A:提取核心关键词
对于短文本,提取关键词往往比生成摘要更实用,Gensim自带这个功能:
from gensim.summarization import keywords text = "short paragraph about natural language processing and Gensim library" print('Keywords:') print(keywords(text))
思路B:使用专门处理短文本的摘要库
推荐用sumy,它支持多种摘要算法,对短文本兼容性更好。先安装库:
pip install sumy
然后用这段代码实现短文本摘要:
from sumy.parsers.plaintext import PlaintextParser from sumy.nlp.tokenizers import Tokenizer from sumy.summarizers.lsa import LsaSummarizer text = "short paragraph about natural language processing and Gensim library" parser = PlaintextParser.from_string(text, Tokenizer("english")) summarizer = LsaSummarizer() # 设置要提取的句子数,这里设为1 summary = summarizer(parser.document, sentences_count=1) print('Summary:') for sentence in summary: print(sentence)
内容的提问来源于stack exchange,提问作者Thankless
相关产品推荐
相关产品推荐

