You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历Pandas DataFrame并对文本进行正负情感分类?

给分词后推文添加情感标签的实现方案

Got it,我来一步步帮你完成给分词推文标注情感标签的操作,完全用循环遍历的方式实现:

步骤1:准备数据与基础环境

先把你给出的示例数据转换成Pandas DataFrame,同时导入依赖库:

import pandas as pd

# 你的示例分词推文数据
tokenized_tweets = ['football, was, good, we, played, well' , 'We, were, unlucky, today, bad, luck' , 'terrible, performance, bad, game']

# 转换为DataFrame
df = pd.DataFrame({'tokenized_tweets': tokenized_tweets})

步骤2:定义情感词库

我们先自定义一组正面和负面的情感关键词,你可以根据推文的领域(比如这里是体育)随时扩展这个词库:

# 自定义情感词库(可按需扩展)
positive_words = {'good', 'well', 'great', 'excellent', 'win', 'perfect'}
negative_words = {'bad', 'unlucky', 'terrible', 'awful', 'lose', 'poor'}

步骤3:循环遍历标注情感

接下来我们逐行遍历DataFrame,处理每条分词推文,通过统计正负词数量来判断情感:

# 初始化空列表存储情感标签
sentiment_labels = []

# 循环遍历每条推文
for tweet in df['tokenized_tweets']:
    # 把逗号分隔的字符串拆成单个词,同时去空格、转小写(避免大小写干扰判断)
    words = [word.strip().lower() for word in tweet.split(',')]
    
    # 统计正负词出现次数
    pos_count = sum(1 for word in words if word in positive_words)
    neg_count = sum(1 for word in words if word in negative_words)
    
    # 判断并添加标签
    if pos_count > neg_count:
        sentiment_labels.append('positive')
    elif neg_count > pos_count:
        sentiment_labels.append('negative')
    else:
        # 正负词数量相等或都没有时,可自定义默认标签,这里暂设为negative
        sentiment_labels.append('negative')

# 将标签列加入原DataFrame
df['sentiment'] = sentiment_labels

查看最终结果

运行完代码后,打印DataFrame就能看到标注好的结果:

print(df)

输出示例:

tokenized_tweets sentiment
0  football, was, good, we, played, well   positive
1        We, were, unlucky, today, bad, luck   negative
2             terrible, performance, bad, game   negative

小优化建议

  • 可以针对体育领域扩展情感词库,比如加入win、score这类正面词,lose、miss这类负面词
  • 如果分词结果包含时态变化(比如played),可以加入词形还原处理,让判断更准确
  • 要是需要更精准的情感分析,后续也可以替换成预训练模型(比如VADER),但如果必须用循环遍历,上面的方法完全满足需求

内容的提问来源于stack exchange,提问作者Joe Pearson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:23:04