如何基于共享前缀对成对列水平排序并同步关联数据?
解决方法
我们可以通过对比原数据的节点列与排序后的节点列,判断是否需要同步调换p1和p2的值,具体步骤如下:
- 对
node1、node2列进行排序,得到标准化后的节点列 - 生成布尔标识,标记原
node1与排序后node1是否一致——不一致的行需要交换p1和p2 - 根据标识批量调整
p1、p2的值 - 合并排序后的节点列与调整后的p列,得到最终结果
完整代码
import pandas as pd import numpy as np # 原数据 df = pd.DataFrame({'node1': ['abc-1', 'xyz-1', 'abc-1', 'xyz-2', 'xyz-2', 'ghi-2'], 'p1': [1, 10, 3, 1, 2, 6], 'p2': [9, 2, 11, 4, 5, 3], 'node2': ['xyz-1', 'abc-1', 'xyz-1', 'def-2', 'def-2', 'xyz-1']}) # 对节点列排序 sorted_nodes = df[['node1', 'node2']].apply(np.sort, axis=1).apply(pd.Series) sorted_nodes.columns = ['node1', 'node2'] # 判断是否需要交换p值:原node1与排序后node1不等时触发交换 swap_flag = df['node1'] != sorted_nodes['node1'] # 批量处理p列:需要交换的行取原p2、p1,不需要的保持原样 p1 = np.where(swap_flag, df['p2'], df['p1']) p2 = np.where(swap_flag, df['p1'], df['p2']) # 合并所有列生成最终结果 final_df = pd.DataFrame({ 'node1': sorted_nodes['node1'], 'p1': p1, 'p2': p2, 'node2': sorted_nodes['node2'] }) print(final_df)
输出结果
node1 p1 p2 node2 0 abc-1 1 9 xyz-1 1 abc-1 2 10 xyz-1 2 abc-1 3 11 xyz-1 3 def-2 4 1 xyz-2 4 def-2 5 2 xyz-2 5 ghi-2 6 3 xyz-1
关键逻辑说明
swap_flag作为批量判断的标识,精准定位需要调换p值的行,避免逐行循环的低效操作np.where实现向量式的条件赋值,高效完成p列的同步调换
内容的提问来源于stack exchange,提问作者VERBOSE
相关产品推荐
相关产品推荐

