You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas层级结构DataFrame转换为指定格式的字典?

解决DataFrame转母子映射字典的问题

我来帮你修正这个问题!你的代码逻辑问题在于每次找到一个mother时,会把DataFrame里所有的child都塞进列表里,没有区分哪些child属于当前的mother。咱们换个更高效的思路来实现:

正确步骤解析

核心思路是先给每个child行标记对应的mother,再分组聚合:

  1. 用ffill()向前填充A列的缺失值,让每个child行都能关联到最近的mother;
  2. 过滤掉B列为空的行(也就是mother本身的行);
  3. 按A列分组,把每个mother对应的child收集成列表,最后转成字典。

完整代码实现

import pandas as pd
import numpy as np

# 构造示例DataFrame
data = {'A': ['mother1', np.nan, np.nan, 'mother2', np.nan, 'mother3', np.nan, np.nan, np.nan],
        'B': [np.nan, 'child1', 'child2', np.nan, 'child1', np.nan, 'child1', 'child2', 'child3']}
df = pd.DataFrame(data)

# 步骤1:向前填充A列,给每个child分配对应的mother
df['A'] = df['A'].ffill()

# 步骤2:过滤掉B列为空的行(保留child行)
filtered_df = df.dropna(subset=['B'])

# 步骤3:分组聚合转字典
results = filtered_df.groupby('A')['B'].apply(list).to_dict()

print(results)

输出结果

{'mother1': ['child1', 'child2'], 'mother2': ['child1'], 'mother3': ['child1', 'child2', 'child3']}

为什么你的代码出错?

你的内层循环遍历了整个DataFrame的所有行,只要B列不是NaN就添加到children列表,这导致每个mother都绑定了所有child的集合,而不是属于它自己的那部分child。用ffill()分组的方式不仅更简洁,效率也比嵌套循环高得多(尤其是数据量大的时候)。

内容的提问来源于stack exchange,提问作者FaCoffee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:34:01