You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:基于字典父子关系规则填充DataFrame空行

解决DataFrame中空值的递归填充问题

需求说明

需要填充以下DataFrame中的空行:

import pandas as pd
import numpy as np

df = pd.DataFrame({ 
    'id':['A', 'B', 'C','D','J','K','Z','Y','H','G'], 
    'test1':[10, 9, 8,7,np.nan,6,np.nan,np.nan,5,np.nan]
})

填充规则基于以下字典(定义每个id的子节点及子节点数量):

dic1={
    'A': [['K', 'J'], 2.0],
    'B': [np.nan, np.nan],
    'C': [['Y'], 1.0],
    'D': [['B', 'C'], 2.0],
    'J': [np.nan, np.nan],
    'K': [np.nan, np.nan],
    'G': [['A', 'H'], 2.0],
    'Y': [['Z'], 1.0],
    'H': [np.nan, np.nan],
    'Z': [['G'], 1.0]
}

空值对应的id列表为:new_list=['J', 'G', 'Y', 'Z'],需遵循以下填充规则:

  • 规则1:若id在字典中无子女(值为NaN),则test1赋值为0;
  • 规则2:若id有1个非new_list中的子女,则test1填充为该子女的test1值;
  • 规则3:若id有多个非new_list中的子女,则test1填充为所有子女test1值的最大值;
  • 规则4:若id的子女属于new_list,则需先递归按规则1-3填充该子女的test1,再填充当前id的test1。

期望输出

df = pd.DataFrame({ 
    'id':['A', 'B', 'C','D','J','K','Z','Y','H','G'], 
    'test1':[10, 9, 8,7,0,6,10,10,5,10]
})

现有代码问题

已实现规则1-3,但无法处理规则4的递归依赖逻辑,现有代码如下:

new_list=['J', 'G', 'Y', 'Z']
dic_df=dict(zip(df.task_id, df.test1))  # 存在笔误:应为df['id']而非df.task_id
act_aa={}
def test_newcase(i):
    if str(dic1[i][0])=='nan':
        df.loc[df['task_id'] == i, ['test1']] = 0  # 存在笔误:应为df['id']而非df.task_id
    else:  
        if any(x not in new_list for x in dic1[i][0]):
            if dic1[i][1]==1.0:
                for k in dic1[i][0]:
                    df.loc[df['task_id'] == i, ['test1','test2']] = dic_df[k][0]  # test2未定义,属于笔误
            else:
                for k in dic1[i][0]:
                    act_aa[k]=str(dic_df[k])[0]
                if act_aa :
                    df.loc[df['task_id'] == i, ['test1']] = max(act_aa.values())      
for i in new_list:
    test_newcase(i)
df

解决方案

通过递归函数处理嵌套依赖,修改后的代码如下:

import pandas as pd
import numpy as np

# 初始化DataFrame
df = pd.DataFrame({ 
    'id':['A', 'B', 'C','D','J','K','Z','Y','H','G'], 
    'test1':[10, 9, 8,7,np.nan,6,np.nan,np.nan,5,np.nan]
})

# 依赖字典
dic1={
    'A': [['K', 'J'], 2.0],
    'B': [np.nan, np.nan],
    'C': [['Y'], 1.0],
    'D': [['B', 'C'], 2.0],
    'J': [np.nan, np.nan],
    'K': [np.nan, np.nan],
    'G': [['A', 'H'], 2.0],
    'Y': [['Z'], 1.0],
    'H': [np.nan, np.nan],
    'Z': [['G'], 1.0]
}

new_list = ['J', 'G', 'Y', 'Z']
# 将test1转为字典,方便递归过程中快速访问和更新
test1_dict = df.set_index('id')['test1'].to_dict()

def fill_test1(id_val):
    # 若当前id的test1已有有效值,直接返回
    if not pd.isna(test1_dict[id_val]):
        return test1_dict[id_val]
    
    children_info = dic1[id_val]
    children = children_info[0]
    
    # 规则1:无子女,赋值为0
    if pd.isna(children):
        test1_dict[id_val] = 0.0
        return 0.0
    
    child_values = []
    # 先递归处理所有属于new_list的子女,确保子女值已填充
    for child in children:
        if child in new_list:
            fill_test1(child)
        # 获取子女的test1值(已填充或原始有效值)
        child_val = test1_dict[child]
        if not pd.isna(child_val):
            child_values.append(child_val)
    
    # 规则2:单个有效子女值,直接赋值
    if len(child_values) == 1:
        test1_dict[id_val] = child_values[0]
    # 规则3:多个有效子女值,取最大值
    elif len(child_values) > 1:
        test1_dict[id_val] = max(child_values)
    
    return test1_dict[id_val]

# 遍历需要填充的id,执行递归填充
for id_val in new_list:
    fill_test1(id_val)

# 将更新后的值同步回DataFrame
df['test1'] = df['id'].map(test1_dict)
print(df)

关键修改点

  1. 使用test1_dict存储test1值,避免频繁操作DataFrame,提升效率同时简化递归中的值访问与更新;
  2. 递归函数fill_test1优先处理当前id的子女(若属于new_list则递归填充),确保依赖的子女值先被计算;
  3. 严格遵循4条填充规则,处理完所有子女后再计算当前id的test1值;
  4. 最后将字典中的更新值同步回DataFrame,保证结果正确。

内容的提问来源于stack exchange,提问作者bb_zz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 06:05:21