如何按字典键结构合并pandas DataFrame的列?
Pandas:将带层级结构的列名转换为嵌套字典/列表
问题
给定带层级命名的列(如B.C、E[0].G[0])的DataFrame,需要将这些列转换为嵌套的字典或列表结构,仅保留顶层键作为最终列,实现从扁平列到嵌套结构的转换。
初始DataFrame代码:
import pandas as pd d = [{ "A" : 1, "B.C" : 2, "B.D" : 3, "E[0].F" : 4, "E[0].G[0]" : 5, }, { "A" : 6, "B.C" : 7, "B.D" : 8, "E[0].F" : 9, "E[0].G[0]" : 10, }] df = pd.DataFrame(d)
初始输出:
A B.C B.D E[0].F E[0].G[0] 0 1 2 3 4 5 1 6 7 8 9 10
期望结果:
A B E 0 1 {'C': 2, 'D': 3} [{'F': 4, 'G': [5]}] 1 6 {'C': 7, 'D': 8} [{'F': 9, 'G': [10]}]
解决方案
通过自定义函数解析列名的层级规则(.表示字典键,[]表示列表索引),逐行构建嵌套结构,再生成新的DataFrame。
1. 编写层级解析函数
这个函数负责将单个列名和对应值,插入到对应的嵌套结构中:
def build_nested(col_name, value): parts = col_name.split('.') root = {} current = root for i, part in enumerate(parts): # 处理列表项,如E[0] if '[' in part and ']' in part: key, idx_str = part.split('[') idx = int(idx_str.rstrip(']')) # 初始化列表(如果不存在) if key not in current: current[key] = [] # 扩展列表到目标索引长度,确保索引有效 while len(current[key]) <= idx: current[key].append({}) # 如果是最后一个部分,直接赋值;否则进入该列表项 if i == len(parts) - 1: current[key][idx] = value else: current = current[key][idx] else: # 处理字典键 if i == len(parts) - 1: current[part] = value else: if part not in current: current[part] = {} current = current[part] # 返回顶层键和对应的嵌套结构 return parts[0], root[parts[0]]
2. 处理整行数据,合并嵌套结构
遍历每一行,将同一顶层键的所有嵌套结构合并,生成新的行数据:
def convert_to_nested(df): new_rows = [] for _, row in df.iterrows(): row_dict = {} for col, val in row.items(): top_key, nested_val = build_nested(col, val) if top_key not in row_dict: row_dict[top_key] = nested_val else: # 合并同一顶层键下的结构 if isinstance(row_dict[top_key], dict) and isinstance(nested_val, dict): row_dict[top_key].update(nested_val) elif isinstance(row_dict[top_key], list) and isinstance(nested_val, list): # 合并列表中的字典项 for idx in range(max(len(row_dict[top_key]), len(nested_val))): if idx >= len(row_dict[top_key]): row_dict[top_key].append(nested_val[idx]) elif isinstance(row_dict[top_key][idx], dict) and isinstance(nested_val[idx], dict): row_dict[top_key][idx].update(nested_val[idx]) new_rows.append(row_dict) return pd.DataFrame(new_rows)
3. 运行测试
result_df = convert_to_nested(df) print(result_df)
输出结果与期望一致:
A B E 0 1 {'C': 2, 'D': 3} [{'F': 4, 'G': [5]}] 1 6 {'C': 7, 'D': 8} [{'F': 9, 'G': [10]}]
适用场景
该方案支持任意深度的嵌套结构,包括字典与列表混合的复杂列名(如E[0].G[1].H.I[3]),能自动处理同一顶层键下的子结构合并。
内容的提问来源于stack exchange,提问作者airline33
相关产品推荐
相关产品推荐

