You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按字典键结构合并pandas DataFrame的列?

Pandas:将带层级结构的列名转换为嵌套字典/列表

问题

给定带层级命名的列(如B.C、E[0].G[0])的DataFrame,需要将这些列转换为嵌套的字典或列表结构,仅保留顶层键作为最终列,实现从扁平列到嵌套结构的转换。

初始DataFrame代码:

import pandas as pd

d = [{
"A" : 1,
"B.C" : 2,
"B.D" : 3,
"E[0].F" : 4,
"E[0].G[0]" : 5,
},
{
"A" : 6,
"B.C" : 7,
"B.D" : 8,
"E[0].F" : 9,
"E[0].G[0]" : 10,
}]

df = pd.DataFrame(d)

初始输出:

A  B.C  B.D  E[0].F  E[0].G[0]
0  1    2    3       4          5
1  6    7    8       9         10

期望结果:

A                   B                     E
0  1  {'C': 2, 'D': 3}  [{'F': 4, 'G': [5]}]
1  6  {'C': 7, 'D': 8}  [{'F': 9, 'G': [10]}]

解决方案

通过自定义函数解析列名的层级规则(.表示字典键,[]表示列表索引),逐行构建嵌套结构,再生成新的DataFrame。

1. 编写层级解析函数

这个函数负责将单个列名和对应值,插入到对应的嵌套结构中:

def build_nested(col_name, value):
    parts = col_name.split('.')
    root = {}
    current = root

    for i, part in enumerate(parts):
        # 处理列表项,如E[0]
        if '[' in part and ']' in part:
            key, idx_str = part.split('[')
            idx = int(idx_str.rstrip(']'))
            # 初始化列表(如果不存在)
            if key not in current:
                current[key] = []
            # 扩展列表到目标索引长度,确保索引有效
            while len(current[key]) <= idx:
                current[key].append({})
            # 如果是最后一个部分,直接赋值;否则进入该列表项
            if i == len(parts) - 1:
                current[key][idx] = value
            else:
                current = current[key][idx]
        else:
            # 处理字典键
            if i == len(parts) - 1:
                current[part] = value
            else:
                if part not in current:
                    current[part] = {}
                current = current[part]
    
    # 返回顶层键和对应的嵌套结构
    return parts[0], root[parts[0]]

2. 处理整行数据,合并嵌套结构

遍历每一行,将同一顶层键的所有嵌套结构合并,生成新的行数据:

def convert_to_nested(df):
    new_rows = []
    for _, row in df.iterrows():
        row_dict = {}
        for col, val in row.items():
            top_key, nested_val = build_nested(col, val)
            if top_key not in row_dict:
                row_dict[top_key] = nested_val
            else:
                # 合并同一顶层键下的结构
                if isinstance(row_dict[top_key], dict) and isinstance(nested_val, dict):
                    row_dict[top_key].update(nested_val)
                elif isinstance(row_dict[top_key], list) and isinstance(nested_val, list):
                    # 合并列表中的字典项
                    for idx in range(max(len(row_dict[top_key]), len(nested_val))):
                        if idx >= len(row_dict[top_key]):
                            row_dict[top_key].append(nested_val[idx])
                        elif isinstance(row_dict[top_key][idx], dict) and isinstance(nested_val[idx], dict):
                            row_dict[top_key][idx].update(nested_val[idx])
        new_rows.append(row_dict)
    return pd.DataFrame(new_rows)

3. 运行测试

result_df = convert_to_nested(df)
print(result_df)

输出结果与期望一致:

A                   B                     E
0  1  {'C': 2, 'D': 3}  [{'F': 4, 'G': [5]}]
1  6  {'C': 7, 'D': 8}  [{'F': 9, 'G': [10]}]

适用场景

该方案支持任意深度的嵌套结构,包括字典与列表混合的复杂列名(如E[0].G[1].H.I[3]),能自动处理同一顶层键下的子结构合并。


内容的提问来源于stack exchange,提问作者airline33

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 12:05:33