You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按ID统计字典中情感词的去重频次

按ID分组统计词汇类别频次(同一ID同类别仅计1次)

输入数据

import pandas as pd
import re  # 需导入正则模块

data = [['This is a long sentence which contains a lot of words among them happy', 1],
       ['This is another sentence which contains the word happy* with special character', 1],
       ['Content and merry are another words which implies happy', 2],
       ['Sad is not happy', 2],
       ['unfortunate has negative conotations', 1]]
df = pd.DataFrame(data, columns=['string', 'id'])
words = {
    "positive" : ["happy", "content"],
    "negative" : ["sad", "unfortunate"],
    "neutral" : ["neutral", "000"]
    }

需求

  • 按id分组
  • 每组内检查是否包含某类别下至少一个词汇,是则该类别计1次(同一ID同类别重复出现只算1次)
  • 汇总所有分组的计数结果,得到各词汇类别的总频次

预期输出:

word  freq
0  positive     2
1  negative     2
2   neutral     0

实现代码

# 为每个类别添加匹配标记列,标记单条记录是否包含该类别词汇
for cat, terms in words.items():
    # 构建正则模式:匹配独立单词、不区分大小写、转义特殊字符
    regex_pattern = r'\b(' + '|'.join(re.escape(t.lower()) for t in terms) + r')\b'
    df[cat] = df['string'].str.lower().str.contains(regex_pattern).astype(int)

# 按id分组,每个类别取最大值(组内只要有一次匹配就计1)
grouped_counts = df.groupby('id')[words.keys()].max()

# 汇总各类别总频次并整理成目标格式
result = grouped_counts.sum().reset_index()
result.columns = ['word', 'freq']

print(result)

关键说明

  • re.escape():处理词汇中的特殊字符(如happy*里的*),避免正则语法错误
  • str.lower():统一转为小写,实现不区分大小写的匹配
  • \b:确保匹配独立单词,防止类似unhappy被误判为包含happy
  • groupby().max():保证同一ID下同一类别无论匹配多少次都只计1次,最后sum()得到所有ID的总次数

内容的提问来源于stack exchange,提问作者Slartibartfast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 16:57:30