如何使用Python将带情感得分的枚举文档合并至原始文档格式
用Python实现文档情感得分计算与合并流程
没问题,我来一步步帮你搞定这个需求!咱们分三个核心步骤来实现,每一步都给你具体的代码和解释:
1. 拆分文档为句子
首先得把包含多句的文档拆成单个句子,这里最方便的工具是nltk的sent_tokenize函数,它能准确识别句子边界。
先安装nltk并下载必要的模型:
import nltk nltk.download('punkt')
拆分句子的代码示例:
from nltk.tokenize import sent_tokenize # 假设你的原始文档列表是这样的 documents = [ "I love this product! It's so easy to use and works perfectly.", "The service was terrible. I waited 2 hours and no one helped me.", "Just okay, not great but not bad either." ] # 拆分每个文档为句子列表 tokenized_docs = [sent_tokenize(doc) for doc in documents]
2. 计算每个句子的情感得分
推荐用VADER(Valence Aware Dictionary and sEntiment Reasoner)来计算情感得分,它专门针对社交媒体和日常文本,对短句的情感判断很准确,而且不需要额外训练。
先安装vaderSentiment:
pip install vaderSentiment
计算得分的代码:
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer # 初始化情感分析器 analyzer = SentimentIntensityAnalyzer() # 遍历每个拆分后的文档,计算每个句子的得分 sentiment_scores = [] for doc in tokenized_docs: doc_scores = [] for sentence in doc: # VADER返回的是一个字典,其中'compound'是综合情感得分,范围从-1(极负)到1(极正) score = analyzer.polarity_scores(sentence)['compound'] doc_scores.append(score) sentiment_scores.append(doc_scores)
3. 合并回原始文档并计算平均分
最后一步就是把每个文档的所有句子得分取平均,和原始文档对应起来:
# 计算每个文档的平均情感得分 average_scores = [sum(scores)/len(scores) if scores else 0 for scores in sentiment_scores] # 合并成原始文档+对应平均分的结果 result = list(zip(documents, average_scores)) # 打印结果看看 for doc, avg_score in result: print(f"文档: {doc}") print(f"平均情感得分: {avg_score:.4f}\n")
完整示例代码
把上面的步骤整合起来,完整代码如下:
import nltk nltk.download('punkt') from nltk.tokenize import sent_tokenize from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer # 原始文档列表 documents = [ "I love this product! It's so easy to use and works perfectly.", "The service was terrible. I waited 2 hours and no one helped me.", "Just okay, not great but not bad either.", "Single sentence document." # 测试单句文档的情况 ] # 步骤1:拆分句子 tokenized_docs = [sent_tokenize(doc) for doc in documents] # 步骤2:初始化情感分析器并计算每句得分 analyzer = SentimentIntensityAnalyzer() sentiment_scores = [] for doc in tokenized_docs: doc_scores = [] for sentence in doc: score = analyzer.polarity_scores(sentence)['compound'] doc_scores.append(score) sentiment_scores.append(doc_scores) # 步骤3:计算平均分并合并结果 average_scores = [sum(scores)/len(scores) if scores else 0 for scores in sentiment_scores] result = list(zip(documents, average_scores)) # 输出结果 for idx, (doc, avg_score) in enumerate(result, 1): print(f"文档 {idx}: {doc}") print(f"平均情感得分: {avg_score:.4f}\n")
补充说明
- 如果你的文档是更正式的书面语,也可以考虑用
TextBlob或者spaCy的情感分析模型,不过VADER对于日常文本的表现已经很出色了。 - 处理单句文档时,代码里的判断
if scores else 0会直接返回该句子的得分(因为len(scores)=1,sum/scores就是它本身),不用额外调整。
内容的提问来源于stack exchange,提问作者Ashok Kumar Jayaraman
相关产品推荐
相关产品推荐

