如何用Pandas自动化统计对应单元格中指定单词的出现频率
Pandas 实现单元格级别的指定词频统计
你可以通过apply()方法逐行处理DataFrame,实现每行对应单元格的词频统计,替代手动生成count列的操作。
基础实现方案
直接利用字符串的count()方法,结合apply()逐行传入当前行的Items和Specified_Word进行统计:
import pandas as pd # 初始化数据 data = {'Company': ['Nike', 'Levi', 'Dell'], 'Items': ['Running Shoes, Walking Shoes, Socks', 'Jeans, Jackets, Designer Shoes', 'Laptops'], 'Specified_Word':['Shoes', 'Shoes', 'Laptops']} df = pd.DataFrame(data) # 自动生成count列:逐行统计Specified_Word在Items中的出现次数 df['count'] = df.apply(lambda row: row['Items'].count(row['Specified_Word']), axis=1) print(df.head())
执行后输出结果:
Company Items Specified_Word count 0 Nike Running Shoes, Walking Shoes, Socks Shoes 2 1 Levi Jeans, Jackets, Designer Shoes Shoes 1 2 Dell Laptops Laptops 1
精确匹配优化方案
如果需要精确匹配完整单词(避免统计类似Shoesmith这类包含目标词但并非独立单词的情况),可以结合正则表达式的单词边界\b来实现:
import pandas as pd import re data = {'Company': ['Nike', 'Levi', 'Dell'], 'Items': ['Running Shoes, Walking Shoes, Socks', 'Jeans, Jackets, Designer Shoes', 'Laptops'], 'Specified_Word':['Shoes', 'Shoes', 'Laptops']} df = pd.DataFrame(data) # 正则精确匹配完整单词,避免部分匹配 df['count'] = df.apply( lambda row: len(re.findall(r'\b' + re.escape(row['Specified_Word']) + r'\b', row['Items'])), axis=1 ) print(df.head())
这里用re.escape()处理目标词,防止特殊字符干扰正则匹配,确保统计结果更准确。
内容的提问来源于stack exchange,提问作者desert_ranger
相关产品推荐
相关产品推荐

