如何将Pandas层级结构DataFrame转换为指定格式的字典?
解决DataFrame转母子映射字典的问题
我来帮你修正这个问题!你的代码逻辑问题在于每次找到一个mother时,会把DataFrame里所有的child都塞进列表里,没有区分哪些child属于当前的mother。咱们换个更高效的思路来实现:
正确步骤解析
核心思路是先给每个child行标记对应的mother,再分组聚合:
- 用
ffill()向前填充A列的缺失值,让每个child行都能关联到最近的mother; - 过滤掉B列为空的行(也就是mother本身的行);
- 按A列分组,把每个mother对应的child收集成列表,最后转成字典。
完整代码实现
import pandas as pd import numpy as np # 构造示例DataFrame data = {'A': ['mother1', np.nan, np.nan, 'mother2', np.nan, 'mother3', np.nan, np.nan, np.nan], 'B': [np.nan, 'child1', 'child2', np.nan, 'child1', np.nan, 'child1', 'child2', 'child3']} df = pd.DataFrame(data) # 步骤1:向前填充A列,给每个child分配对应的mother df['A'] = df['A'].ffill() # 步骤2:过滤掉B列为空的行(保留child行) filtered_df = df.dropna(subset=['B']) # 步骤3:分组聚合转字典 results = filtered_df.groupby('A')['B'].apply(list).to_dict() print(results)
输出结果
{'mother1': ['child1', 'child2'], 'mother2': ['child1'], 'mother3': ['child1', 'child2', 'child3']}
为什么你的代码出错?
你的内层循环遍历了整个DataFrame的所有行,只要B列不是NaN就添加到children列表,这导致每个mother都绑定了所有child的集合,而不是属于它自己的那部分child。用ffill()分组的方式不仅更简洁,效率也比嵌套循环高得多(尤其是数据量大的时候)。
内容的提问来源于stack exchange,提问作者FaCoffee
相关产品推荐
相关产品推荐

