pandas中如何获取两列整数之间的关联链路分组
解题思路
我们可以通过构造列之间的映射关系,再定位链路起始节点的方式完成关联链路提取,具体步骤如下:
- 过滤掉第二列值为0的行,生成第一列到第二列的映射字典
- 提取所有仅出现在第一列、未出现在第二列的数值作为链路起始节点
- 遍历每个起始节点,顺着映射字典迭代拼接链路,直到无法找到下一个关联值为止
代码实现
首先构造测试用的DataFrame:
import pandas as pd # 构造示例DataFrame,和题目示例匹配 df = pd.DataFrame({ 'col1': [2, 3, 4, 5, 6, 7, 8], 'col2': [3, 4, 5, 0, 7, 8, 0] })
执行链路提取逻辑:
# 生成映射字典,过滤第二列为0的非关联行 mapping = df[df['col2'] != 0].set_index('col1')['col2'].to_dict() # 定位所有链路起始节点:存在于第一列、但不存在于第二列的数值 all_col2_values = set(df['col2']) start_nodes = [val for val in df['col1'] if val not in all_col2_values] # 遍历生成完整链路列表 result = [] for start in start_nodes: current_chain = [start] current_node = start while current_node in mapping: current_node = mapping[current_node] current_chain.append(current_node) result.append(current_chain) print(result)
输出结果
运行代码后得到的结果和预期完全一致:
[[2, 3, 4, 5], [6, 7, 8]]
内容的提问来源于stack exchange,提问作者Samsudeen Bankole
相关产品推荐
相关产品推荐

