You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas两列中的数值关联为成员分组?

解决Pandas两列数值关联分组的问题

咱先把需求掰扯清楚:你想要把DataFrame里A、B两列的数值按照「相互关联」的逻辑分组——比如A列的0对应B列的1,A列的1又对应B列的0,这俩明显是一伙的;再看A列的8对应B列的112、9、114,而A列的9又关联112、8、114,甚至还有134、135,那这些数全得归到同一个组里对吧?这本质就是图论里的连通分量问题,我给你两种实用的实现方案:

方法一:用NetworkX快速实现(简洁易上手)

NetworkX是专门处理图结构的工具库,能直接帮我们找出所有连通的节点集合,适合快速验证需求。

步骤1:先装依赖(没装过的话)

pip install networkx

步骤2:完整代码实现

import pandas as pd
import networkx as nx

# 你的原始DataFrame(我补全了B列末尾的缺失值,你可以根据实际数据调整)
df = pd.DataFrame({
    'A':[0, 1, 3, 4, 6, 7, 8, 8, 8, 9, 9, 9, 9, 9, 11, 12, 13, 14, 15, 15, 15, 16, 16, 16, 16, 17, 17, 17, 17, 18, 18, 18, 18, 18, 19, 19, 19, 19, 20, 20, 21, 22, 24, 25, 26, 27, 28, 29, 29],
    'B':[1, 0, 4, 3, 7, 6, 112, 9, 114, 134, 135, 112, 8, 114, 14, 13, 12, 11, 16, 17, 18, 17, 15, 18, 19, 16, 18, 15, 19, 17, 16, 15, 19, 20, 20, 18, 17, 16, 19, 18, 22, 21, 25, 24, 27, 26, 29, 28, 28]
})

# 1. 构建无向图:把每一行的A和B看作一条连接的边
G = nx.Graph()
edges = list(zip(df['A'], df['B']))
G.add_edges_from(edges)

# 2. 找出所有连通分量,给每个分量分配唯一组ID
component_id = {}
for group_num, component in enumerate(nx.connected_components(G)):
    for num in component:
        component_id[num] = group_num

# 3. 把组ID映射回原DataFrame
df['A_group'] = df['A'].map(component_id)
df['B_group'] = df['B'].map(component_id)
# 也可以直接给每行标记组(因为A和B属于同一组,用A或B的组ID都可以)
df['group_id'] = df['A'].map(component_id)

# 查看前10行结果
print(df[['A', 'B', 'group_id']].head(10))

结果说明

运行后你会看到:

  • 0和1的group_id都是0
  • 8、9、112、114、134、135的group_id都是2
  • 所有互相连通的数值都会共享同一个组ID

方法二:用并查集(Union-Find)实现(高效适合大数据)

如果你的数据集特别大(百万级以上),并查集的效率会比NetworkX高很多,它是专门处理动态连通性问题的数据结构,适合生产环境的大规模数据处理。

完整代码实现

import pandas as pd

# 实现并查集类(带路径压缩优化)
class UnionFind:
    def __init__(self):
        self.parent = {}
    
    def find(self, x):
        # 查找节点x的根节点,路径压缩减少后续查找成本
        if x not in self.parent:
            self.parent[x] = x
        if self.parent[x] != x:
            self.parent[x] = self.find(self.parent[x])
        return self.parent[x]
    
    def union(self, x, y):
        # 合并两个节点所在的集合
        x_root = self.find(x)
        y_root = self.find(y)
        if x_root != y_root:
            self.parent[y_root] = x_root

# 你的原始DataFrame
df = pd.DataFrame({
    'A':[0, 1, 3, 4, 6, 7, 8, 8, 8, 9, 9, 9, 9, 9, 11, 12, 13, 14, 15, 15, 15, 16, 16, 16, 16, 17, 17, 17, 17, 18, 18, 18, 18, 18, 19, 19, 19, 19, 20, 20, 21, 22, 24, 25, 26, 27, 28, 29, 29],
    'B':[1, 0, 4, 3, 7, 6, 112, 9, 114, 134, 135, 112, 8, 114, 14, 13, 12, 11, 16, 17, 18, 17, 15, 18, 19, 16, 18, 15, 19, 17, 16, 15, 19, 20, 20, 18, 17, 16, 19, 18, 22, 21, 25, 24, 27, 26, 29, 28, 28]
})

# 初始化并查集
uf = UnionFind()
# 遍历所有行,合并A和B的数值
for a, b in zip(df['A'], df['B']):
    uf.union(a, b)

# 把根节点映射成连续的组编号(让组ID更规整)
root_to_group = {root: idx for idx, root in enumerate(set(uf.find(node) for node in uf.parent))}
component_id = {node: root_to_group[uf.find(node)] for node in uf.parent}

# 映射回DataFrame
df['group_id'] = df['A'].map(component_id)

# 查看前10行结果
print(df[['A', 'B', 'group_id']].head(10))

结果说明

这个方法和NetworkX的输出完全一致,但处理超大数据集时速度会快很多,适合需要高性能的场景。


内容的提问来源于stack exchange,提问作者Kdog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:35:26