Python pandas如何统计DataFrame逗号分隔字符串的元素出现频次
实现方法
核心思路是先把每行逗号分隔的字符串拆分为独立单个元素(注意去除拆分后元素前后的多余空格),再对所有元素做频次统计,两种常用实现路径如下:
方法1:直接基于原DataFrame操作(推荐,效率更高)
不需要提前把列转成列表,直接用pandas内置的字符串拆分+展开方法即可,注意列名大小写匹配(你的代码里列名是Value首字母大写,提取列表时写的小写value会触发报错):
import pandas as pd # 原数据构建 df = pd.DataFrame({'Value': ['Apple, Banana, Pineapple', 'Apple, Orange', 'Celery, Carrot, Beetroot, Cucumber', 'Watermelon, Apple, Lychee']}) # 拆分、去空格、展开、计数 count_result = ( df['Value'] .str.split(',') # 按逗号拆分每行字符串为列表 .explode() # 把列表里的每个元素展开成单独的行 .str.strip() # 去除元素前后的空格(比如拆分后的" Banana"会处理为"Banana") .value_counts() ) print(count_result)
运行后输出的统计结果如下:
Apple 3 Banana 1 Pineapple 1 Orange 1 Celery 1 Carrot 1 Beetroot 1 Cucumber 1 Watermelon 1 Lychee 1 Name: Value, dtype: int64
如果需要纯文本的元素 频次格式,直接遍历结果打印即可:
for item, cnt in count_result.items(): print(f"{item} {cnt}")
方法2:基于你已经提取的listdf列表实现
如果你已经拿到了listdf列表(注意修正列名大小写问题),可以先把所有拆分后的元素展平成一维列表,再做计数,两种写法都可用:
写法1:转pandas Series复用value_counts
# 注意修正列名大小写,原列名是Value不是value listdf = df["Value"].to_list() # 展平所有元素 all_items = [] for s in listdf: # 拆分后去空格,过滤空字符串(防止原字符串末尾多余逗号产生无效空值) items = [i.strip() for i in s.split(',') if i.strip()] all_items.extend(items) # 转Series后计数 count_result = pd.Series(all_items).value_counts()
写法2:用Python标准库collections.Counter计数
不需要依赖pandas的计数方法:
from collections import Counter count_result = Counter(all_items) # 直接打印即可得到键值对形式的频次结果 print(count_result)
注意:如果原字符串存在末尾多余逗号(比如你示例里第二行
Apple, Orange,末尾多了一个逗号),拆分后会产生空字符串,上面代码里的if i.strip()判断会自动过滤这类无效空值,避免统计出错。
内容的提问来源于stack exchange,提问作者Kusisi Karem
相关产品推荐
相关产品推荐

