已清洗评论列表的NLP文本摘要生成方法咨询(禁用heapq)
嘿,针对你这个提取评论常见内容做摘要的需求,我给你几个靠谱的实现思路,都是避开heapq的哈,毕竟你之前用它没成功~
方法1:基于词频/短语频的核心内容提取
这个方法最直接,用户反复提到的词汇或短语句型,肯定是讨论的重点,适合快速抓取核心诉求:
from collections import Counter from nltk.util import ngrams import nltk nltk.download('punkt') # 把所有评论的分词和2-gram短语合并 all_tokens = [] for text in clean_text_summary: tokens = nltk.word_tokenize(text.lower()) all_tokens.extend(tokens) # 提取双词短语(比如"so bad"、"going to deregister"这类更有意义的组合) bigrams = ngrams(tokens, 2) all_tokens.extend([' '.join(bg) for bg in bigrams]) # 统计频率并取Top N高频内容 freq_counter = Counter(all_tokens) top_10_content = freq_counter.most_common(10) # 生成通顺的摘要 summary = "用户评论中最受关注的内容包括:" + ", ".join([item[0] for item in top_10_content]) print(summary)
方法2:基于主题建模的分主题摘要
如果评论涉及多个不同的投诉/讨论点,用LDA主题建模可以把相似内容聚类,能清晰展示不同核心主题:
from gensim import corpora, models import nltk nltk.download('punkt') # 预处理语料 texts = [nltk.word_tokenize(text.lower()) for text in clean_text_summary] dictionary = corpora.Dictionary(texts) corpus = [dictionary.doc2bow(text) for text in texts] # 训练LDA模型(num_topics可以根据实际评论数量调整,比如3-5个主题) lda_model = models.LdaModel(corpus, num_topics=3, id2word=dictionary, passes=15) # 提取每个主题的关键词生成摘要 summary = "用户讨论的核心主题及关键词:\n" for idx, topic in lda_model.show_topics(formatted=False, num_words=5): summary += f"- 主题{idx+1}:{', '.join([word for word, _ in topic])}\n" print(summary)
方法3:基于TF-IDF的核心句子提取
如果想保留用户的原句风格,用TF-IDF计算句子的代表性,提取最能覆盖整体内容的句子作为摘要:
from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.metrics.pairwise import cosine_similarity import numpy as np # 初始化TF-IDF向量化器 tfidf = TfidfVectorizer() tfidf_matrix = tfidf.fit_transform(clean_text_summary) # 计算每个句子与整个语料的平均相似度(相似度越高,句子越具代表性) sentence_similarities = np.mean(cosine_similarity(tfidf_matrix), axis=1) # 取相似度最高的3个句子作为核心摘要 top_sentence_indices = sentence_similarities.argsort()[-3:][::-1] top_sentences = [clean_text_summary[i] for i in top_sentence_indices] summary = "用户评论的核心内容摘要:\n- " + "\n- ".join(top_sentences) print(summary)
这三个方法各有侧重:方法1适合快速抓高频关键词,方法2适合分主题梳理不同诉求,方法3适合保留原句风格的摘要,你可以根据自己的需求选~
内容的提问来源于stack exchange,提问作者Django0602
相关产品推荐
相关产品推荐

