如何基于Pandas DataFrame列值创建嵌套字典?
问题说明
现有如下Pandas DataFrame:
df1 = pd.DataFrame(data={'col1': [21, 44, 28, 32, 20, 39, 42], 'col2': ['<1', '>2', '>2', '>3', '<1', '>2', '>4'], 'col3': ['yes', 'yes', 'no', 'no', 'yes', 'no', 'yes'], 'col4': [1, 1, 0, 0, 1, 0, 1], 'Group': [0, 2, 1, 1, 0, 1, 2] })
DataFrame内容预览:
col1 col2 col3 col4 Group 0 21 <1 yes 1 0 1 44 >2 yes 1 2 2 28 >2 no 0 1 3 32 >3 no 0 1 4 20 <1 yes 1 0 5 39 >2 no 0 1 6 42 >4 yes 1 2
需要基于Group列的取值0、1、2,生成指定结构的嵌套字典,处理规则:
- 数值列(
col1、col4):按分组计算均值(mean)和标准差(sd) - 字符串列(
col2、col3):提取分组内去重后的唯一值,存为列表
目标输出结构如下:
{0: {'col1': {'mean': 20.5, 'sd': 0.5}, 'col2': ['<1'], 'col3': ['yes'], 'col4': {'mean': 1, 'sd': 0}}, 1: {'col1': {'mean': 33, 'sd': 4.54}, 'col2': ['>2', '>3'], 'col3': ['no'], 'col4': {'mean': 0, 'sd': 0}}, 2: {'col1': {'mean': 43, 'sd': 1}, 'col2': ['>2', '>4'], 'col3': ['yes'], 'col4': {'mean': 1, 'sd': 0}}}
实现代码
核心逻辑是按Group列分组遍历,对不同类型的列执行对应计算后组装字典。注意示例中的标准差为总体标准差(除以样本量n),调用std()时需指定ddof=0才能匹配结果。
import pandas as pd # 初始化原始DataFrame df1 = pd.DataFrame(data={'col1': [21, 44, 28, 32, 20, 39, 42], 'col2': ['<1', '>2', '>2', '>3', '<1', '>2', '>4'], 'col3': ['yes', 'yes', 'no', 'no', 'yes', 'no', 'yes'], 'col4': [1, 1, 0, 0, 1, 0, 1], 'Group': [0, 2, 1, 1, 0, 1, 2] }) # 区分数值列和字符串列 numeric_columns = ['col1', 'col4'] string_columns = ['col2', 'col3'] result_dict = {} # 按Group分组迭代处理 for group_id, group_data in df1.groupby('Group'): group_result = {} # 处理数值列:计算均值、总体标准差,保留2位小数 for col in numeric_columns: mean_val = round(group_data[col].mean(), 2) # ddof=0对应总体标准差,匹配示例计算结果 sd_val = round(group_data[col].std(ddof=0), 2) group_result[col] = {'mean': mean_val, 'sd': sd_val} # 处理字符串列:提取去重值转为列表 for col in string_columns: group_result[col] = group_data[col].unique().tolist() result_dict[group_id] = group_result # 打印输出结果 print(result_dict)
注意事项
- 如果需要计算样本标准差(除以n-1),去掉
std()里的ddof=0参数即可 - 代码中
round(...,2)用于控制输出小数位数,可根据实际精度需求调整 - 字符串列返回的唯一值顺序为数据中首次出现的顺序,和示例展示顺序一致
内容的提问来源于stack exchange,提问作者hanzgs
相关产品推荐
相关产品推荐

