You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于共享前缀对成对列水平排序并同步关联数据?

解决方法

我们可以通过对比原数据的节点列与排序后的节点列,判断是否需要同步调换p1和p2的值,具体步骤如下:

  1. 对node1、node2列进行排序,得到标准化后的节点列
  2. 生成布尔标识,标记原node1与排序后node1是否一致——不一致的行需要交换p1和p2
  3. 根据标识批量调整p1、p2的值
  4. 合并排序后的节点列与调整后的p列,得到最终结果

完整代码

import pandas as pd
import numpy as np

# 原数据
df = pd.DataFrame({'node1': ['abc-1', 'xyz-1', 'abc-1', 'xyz-2', 'xyz-2', 'ghi-2'],
 'p1': [1, 10, 3, 1, 2, 6],
 'p2': [9, 2, 11, 4, 5, 3],
 'node2': ['xyz-1', 'abc-1', 'xyz-1', 'def-2', 'def-2', 'xyz-1']})

# 对节点列排序
sorted_nodes = df[['node1', 'node2']].apply(np.sort, axis=1).apply(pd.Series)
sorted_nodes.columns = ['node1', 'node2']

# 判断是否需要交换p值:原node1与排序后node1不等时触发交换
swap_flag = df['node1'] != sorted_nodes['node1']

# 批量处理p列:需要交换的行取原p2、p1,不需要的保持原样
p1 = np.where(swap_flag, df['p2'], df['p1'])
p2 = np.where(swap_flag, df['p1'], df['p2'])

# 合并所有列生成最终结果
final_df = pd.DataFrame({
    'node1': sorted_nodes['node1'],
    'p1': p1,
    'p2': p2,
    'node2': sorted_nodes['node2']
})

print(final_df)

输出结果

node1  p1  p2  node2
0  abc-1   1   9  xyz-1
1  abc-1   2  10  xyz-1
2  abc-1   3  11  xyz-1
3  def-2   4   1  xyz-2
4  def-2   5   2  xyz-2
5  ghi-2   6   3  xyz-1

关键逻辑说明

  • swap_flag 作为批量判断的标识,精准定位需要调换p值的行,避免逐行循环的低效操作
  • np.where 实现向量式的条件赋值,高效完成p列的同步调换

内容的提问来源于stack exchange,提问作者VERBOSE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 19:37:37