You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python将带情感得分的枚举文档合并至原始文档格式

用Python实现文档情感得分计算与合并流程

没问题,我来一步步帮你搞定这个需求!咱们分三个核心步骤来实现,每一步都给你具体的代码和解释:

1. 拆分文档为句子

首先得把包含多句的文档拆成单个句子,这里最方便的工具是nltk的sent_tokenize函数,它能准确识别句子边界。

先安装nltk并下载必要的模型:

import nltk
nltk.download('punkt')

拆分句子的代码示例:

from nltk.tokenize import sent_tokenize

# 假设你的原始文档列表是这样的
documents = [
    "I love this product! It's so easy to use and works perfectly.",
    "The service was terrible. I waited 2 hours and no one helped me.",
    "Just okay, not great but not bad either."
]

# 拆分每个文档为句子列表
tokenized_docs = [sent_tokenize(doc) for doc in documents]

2. 计算每个句子的情感得分

推荐用VADER(Valence Aware Dictionary and sEntiment Reasoner)来计算情感得分,它专门针对社交媒体和日常文本,对短句的情感判断很准确,而且不需要额外训练。

先安装vaderSentiment:

pip install vaderSentiment

计算得分的代码:

from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer

# 初始化情感分析器
analyzer = SentimentIntensityAnalyzer()

# 遍历每个拆分后的文档,计算每个句子的得分
sentiment_scores = []
for doc in tokenized_docs:
    doc_scores = []
    for sentence in doc:
        # VADER返回的是一个字典,其中'compound'是综合情感得分,范围从-1(极负)到1(极正)
        score = analyzer.polarity_scores(sentence)['compound']
        doc_scores.append(score)
    sentiment_scores.append(doc_scores)

3. 合并回原始文档并计算平均分

最后一步就是把每个文档的所有句子得分取平均,和原始文档对应起来:

# 计算每个文档的平均情感得分
average_scores = [sum(scores)/len(scores) if scores else 0 for scores in sentiment_scores]

# 合并成原始文档+对应平均分的结果
result = list(zip(documents, average_scores))

# 打印结果看看
for doc, avg_score in result:
    print(f"文档: {doc}")
    print(f"平均情感得分: {avg_score:.4f}\n")

完整示例代码

把上面的步骤整合起来,完整代码如下:

import nltk
nltk.download('punkt')
from nltk.tokenize import sent_tokenize
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer

# 原始文档列表
documents = [
    "I love this product! It's so easy to use and works perfectly.",
    "The service was terrible. I waited 2 hours and no one helped me.",
    "Just okay, not great but not bad either.",
    "Single sentence document."  # 测试单句文档的情况
]

# 步骤1:拆分句子
tokenized_docs = [sent_tokenize(doc) for doc in documents]

# 步骤2:初始化情感分析器并计算每句得分
analyzer = SentimentIntensityAnalyzer()
sentiment_scores = []
for doc in tokenized_docs:
    doc_scores = []
    for sentence in doc:
        score = analyzer.polarity_scores(sentence)['compound']
        doc_scores.append(score)
    sentiment_scores.append(doc_scores)

# 步骤3:计算平均分并合并结果
average_scores = [sum(scores)/len(scores) if scores else 0 for scores in sentiment_scores]
result = list(zip(documents, average_scores))

# 输出结果
for idx, (doc, avg_score) in enumerate(result, 1):
    print(f"文档 {idx}: {doc}")
    print(f"平均情感得分: {avg_score:.4f}\n")

补充说明

  • 如果你的文档是更正式的书面语,也可以考虑用TextBlob或者spaCy的情感分析模型,不过VADER对于日常文本的表现已经很出色了。
  • 处理单句文档时,代码里的判断if scores else 0会直接返回该句子的得分(因为len(scores)=1,sum/scores就是它本身),不用额外调整。

内容的提问来源于stack exchange,提问作者Ashok Kumar Jayaraman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:38:59