如何使用apply遍历pandas DataFrame并访问下一行值,生成parent父级列
Pandas DataFrame 生成父级字段方案
你的需求核心是每个深度为d的组件,父级为当前行上方最近的深度为d-1的组件,两种实现方案如下:
方案1:遍历实现(效率更高,推荐)
不需要用apply,直接按行遍历存储每个深度的最近组件,适配大表场景:
import pandas as pd # 原始数据 x = {'depth':[1,1,2,2,2,3,3,3,3],"Component":["A","B","C","D","E","F","G","H","I"]} x = pd.DataFrame(x) last_depth_comp = {} parent_col = [] for _, row in x.iterrows(): cur_depth = row['depth'] # 根节点(深度1)无父级 if cur_depth == 1: parent_col.append("None") else: # 取深度减1对应的最近组件作为父级 parent_col.append(last_depth_comp[cur_depth - 1]) # 更新当前深度的最近组件 last_depth_comp[cur_depth] = row['Component'] # 写入父级列,此时x即为目标DataFrame y x['parent'] = parent_col
方案2:apply 实现(适配需要嵌入自定义逻辑的场景)
如果你的业务逻辑必须放在apply中执行,可以借助外部变量存储上下文信息:
last_depth_comp = {} def fetch_parent(row): global last_depth_comp cur_depth = row['depth'] parent = "None" if cur_depth == 1 else last_depth_comp[cur_depth - 1] last_depth_comp[cur_depth] = row['Component'] return parent x['parent'] = x.apply(fetch_parent, axis=1)
注意事项
两种方案均不依赖组件的命名规则,仅通过行顺序和depth字段判断父子关系,适配你提到的无简单关联规则的真实业务场景。
内容的提问来源于stack exchange,提问作者SIUV
相关产品推荐
相关产品推荐

