如何统计dataframe每行文本与指定词表的重叠词数量并新增列
实现方案
直接基于pandas的apply方法就能实现需求,以下是可运行的完整代码:
import pandas as pd # 构造示例DataFrame df = pd.DataFrame({ 'feelings': [ ['happy', 'happy', 'sad'], ['neutral', 'sad', 'mad'], ['neutral', 'neutral', 'happy'] ] }, index=[1,2,3]) # 定义自定义词表,转集合可提升匹配效率 lst1 = {'happy', 'fantastic'} lst2 = {'mad', 'sad'} lst3 = {'neutral'} # 逐行统计各词表的单词出现总次数 df['occlst1'] = df['feelings'].apply(lambda x: sum(word in lst1 for word in x)) df['occlst2'] = df['feelings'].apply(lambda x: sum(word in lst2 for word in x)) df['occlst3'] = df['feelings'].apply(lambda x: sum(word in lst3 for word in x)) print(df)
如果你处理的数据量大、词表数量多,可以先统计每行的词频再批量匹配,减少重复遍历次数:
from collections import Counter def calc_counts(row_counter): return ( sum(row_counter.get(w,0) for w in lst1), sum(row_counter.get(w,0) for w in lst2), sum(row_counter.get(w,0) for w in lst3) ) df[['occlst1','occlst2','occlst3']] = df['feelings']\ .apply(lambda x: pd.Series(calc_counts(Counter(x))))
两种方案输出的结果都和你给出的示例完全一致。
内容的提问来源于stack exchange,提问作者Leonie
相关产品推荐
相关产品推荐

