You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计dataframe每行文本与指定词表的重叠词数量并新增列

实现方案

直接基于pandas的apply方法就能实现需求,以下是可运行的完整代码:

import pandas as pd

# 构造示例DataFrame
df = pd.DataFrame({
    'feelings': [
        ['happy', 'happy', 'sad'],
        ['neutral', 'sad', 'mad'],
        ['neutral', 'neutral', 'happy']
    ]
}, index=[1,2,3])

# 定义自定义词表,转集合可提升匹配效率
lst1 = {'happy', 'fantastic'}
lst2 = {'mad', 'sad'}
lst3 = {'neutral'}

# 逐行统计各词表的单词出现总次数
df['occlst1'] = df['feelings'].apply(lambda x: sum(word in lst1 for word in x))
df['occlst2'] = df['feelings'].apply(lambda x: sum(word in lst2 for word in x))
df['occlst3'] = df['feelings'].apply(lambda x: sum(word in lst3 for word in x))

print(df)

如果你处理的数据量大、词表数量多,可以先统计每行的词频再批量匹配,减少重复遍历次数:

from collections import Counter

def calc_counts(row_counter):
    return (
        sum(row_counter.get(w,0) for w in lst1),
        sum(row_counter.get(w,0) for w in lst2),
        sum(row_counter.get(w,0) for w in lst3)
    )

df[['occlst1','occlst2','occlst3']] = df['feelings']\
    .apply(lambda x: pd.Series(calc_counts(Counter(x))))

两种方案输出的结果都和你给出的示例完全一致。

内容的提问来源于stack exchange,提问作者Leonie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 08:18:04