如何在Pandas中按ID统计字典中情感词的去重频次
按ID分组统计词汇类别频次(同一ID同类别仅计1次)
输入数据
import pandas as pd import re # 需导入正则模块 data = [['This is a long sentence which contains a lot of words among them happy', 1], ['This is another sentence which contains the word happy* with special character', 1], ['Content and merry are another words which implies happy', 2], ['Sad is not happy', 2], ['unfortunate has negative conotations', 1]] df = pd.DataFrame(data, columns=['string', 'id']) words = { "positive" : ["happy", "content"], "negative" : ["sad", "unfortunate"], "neutral" : ["neutral", "000"] }
需求
- 按
id分组 - 每组内检查是否包含某类别下至少一个词汇,是则该类别计1次(同一ID同类别重复出现只算1次)
- 汇总所有分组的计数结果,得到各词汇类别的总频次
预期输出:
word freq 0 positive 2 1 negative 2 2 neutral 0
实现代码
# 为每个类别添加匹配标记列,标记单条记录是否包含该类别词汇 for cat, terms in words.items(): # 构建正则模式:匹配独立单词、不区分大小写、转义特殊字符 regex_pattern = r'\b(' + '|'.join(re.escape(t.lower()) for t in terms) + r')\b' df[cat] = df['string'].str.lower().str.contains(regex_pattern).astype(int) # 按id分组,每个类别取最大值(组内只要有一次匹配就计1) grouped_counts = df.groupby('id')[words.keys()].max() # 汇总各类别总频次并整理成目标格式 result = grouped_counts.sum().reset_index() result.columns = ['word', 'freq'] print(result)
关键说明
re.escape():处理词汇中的特殊字符(如happy*里的*),避免正则语法错误str.lower():统一转为小写,实现不区分大小写的匹配\b:确保匹配独立单词,防止类似unhappy被误判为包含happygroupby().max():保证同一ID下同一类别无论匹配多少次都只计1次,最后sum()得到所有ID的总次数
内容的提问来源于stack exchange,提问作者Slartibartfast
相关产品推荐
相关产品推荐

