You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

已清洗评论列表的NLP文本摘要生成方法咨询(禁用heapq)

嘿,针对你这个提取评论常见内容做摘要的需求,我给你几个靠谱的实现思路,都是避开heapq的哈,毕竟你之前用它没成功~

方法1:基于词频/短语频的核心内容提取

这个方法最直接,用户反复提到的词汇或短语句型,肯定是讨论的重点,适合快速抓取核心诉求:

from collections import Counter
from nltk.util import ngrams
import nltk
nltk.download('punkt')

# 把所有评论的分词和2-gram短语合并
all_tokens = []
for text in clean_text_summary:
    tokens = nltk.word_tokenize(text.lower())
    all_tokens.extend(tokens)
    # 提取双词短语(比如"so bad"、"going to deregister"这类更有意义的组合)
    bigrams = ngrams(tokens, 2)
    all_tokens.extend([' '.join(bg) for bg in bigrams])

# 统计频率并取Top N高频内容
freq_counter = Counter(all_tokens)
top_10_content = freq_counter.most_common(10)

# 生成通顺的摘要
summary = "用户评论中最受关注的内容包括:" + ", ".join([item[0] for item in top_10_content])
print(summary)
方法2:基于主题建模的分主题摘要

如果评论涉及多个不同的投诉/讨论点,用LDA主题建模可以把相似内容聚类,能清晰展示不同核心主题:

from gensim import corpora, models
import nltk
nltk.download('punkt')

# 预处理语料
texts = [nltk.word_tokenize(text.lower()) for text in clean_text_summary]
dictionary = corpora.Dictionary(texts)
corpus = [dictionary.doc2bow(text) for text in texts]

# 训练LDA模型(num_topics可以根据实际评论数量调整,比如3-5个主题)
lda_model = models.LdaModel(corpus, num_topics=3, id2word=dictionary, passes=15)

# 提取每个主题的关键词生成摘要
summary = "用户讨论的核心主题及关键词:\n"
for idx, topic in lda_model.show_topics(formatted=False, num_words=5):
    summary += f"- 主题{idx+1}:{', '.join([word for word, _ in topic])}\n"
print(summary)
方法3:基于TF-IDF的核心句子提取

如果想保留用户的原句风格,用TF-IDF计算句子的代表性,提取最能覆盖整体内容的句子作为摘要:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

# 初始化TF-IDF向量化器
tfidf = TfidfVectorizer()
tfidf_matrix = tfidf.fit_transform(clean_text_summary)

# 计算每个句子与整个语料的平均相似度(相似度越高,句子越具代表性)
sentence_similarities = np.mean(cosine_similarity(tfidf_matrix), axis=1)

# 取相似度最高的3个句子作为核心摘要
top_sentence_indices = sentence_similarities.argsort()[-3:][::-1]
top_sentences = [clean_text_summary[i] for i in top_sentence_indices]

summary = "用户评论的核心内容摘要:\n- " + "\n- ".join(top_sentences)
print(summary)

这三个方法各有侧重:方法1适合快速抓高频关键词,方法2适合分主题梳理不同诉求,方法3适合保留原句风格的摘要,你可以根据自己的需求选~

内容的提问来源于stack exchange,提问作者Django0602

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 10:07:30