如何统计值为列表的字典中元素频次并导出为CSV
统计字典内列表元素频次并导出CSV
我通过df.groupby('column1')['column2'].agg(list).to_dict()从DataFrame生成了一个字典,现在需要统计每个键对应列表内各元素的频次,得到嵌套字典形式的结果,同时将结果导出为.csv文件。
示例输入
df_dict = { 'Apples': ['big', 'medium', 'medium', 'medium','big','small'], 'Oranges': ['big', 'medium', 'big'], 'Bananas': ['small', 'small', 'small','small', 'big'], 'Pineapples': ['small', 'big', 'big','big'] }
期望输出
df_dict_counts = { 'Apples': {'big':2, 'medium':3, 'small':1}, 'Oranges': {'big':2, 'medium':1}, 'Bananas': {'small':4, 'big':1}, 'Pineapples': {'small':1, 'big':3} }
解决方法
1. 生成嵌套频次字典
用collections.Counter快速统计列表元素频次,结合字典推导式遍历原字典,生成目标嵌套结构:
from collections import Counter # 生成嵌套频次字典(Counter是dict子类,如需纯字典可转成dict) df_dict_counts = {fruit: dict(Counter(sizes)) for fruit, sizes in df_dict.items()}
2. 导出为CSV文件
CSV需要扁平结构,以下提供两种实现方式:
方式一:使用Python内置csv模块
import csv with open('fruit_size_counts.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['Fruit', 'Size', 'Count']) # 写入表头 # 遍历嵌套字典写入每行数据 for fruit, size_counts in df_dict_counts.items(): for size, count in size_counts.items(): writer.writerow([fruit, size, count])
方式二:使用pandas(适合已有数据处理场景)
import pandas as pd # 将嵌套字典转为扁平DataFrame df = pd.DataFrame.from_dict(df_dict_counts, orient='index').stack().reset_index() df.columns = ['Fruit', 'Size', 'Count'] # 导出CSV df.to_csv('fruit_size_counts.csv', index=False, encoding='utf-8')
内容的提问来源于stack exchange,提问作者drr
相关产品推荐
相关产品推荐

