如何将Pandas两列中的数值关联为成员分组?
解决Pandas两列数值关联分组的问题
咱先把需求掰扯清楚:你想要把DataFrame里A、B两列的数值按照「相互关联」的逻辑分组——比如A列的0对应B列的1,A列的1又对应B列的0,这俩明显是一伙的;再看A列的8对应B列的112、9、114,而A列的9又关联112、8、114,甚至还有134、135,那这些数全得归到同一个组里对吧?这本质就是图论里的连通分量问题,我给你两种实用的实现方案:
方法一:用NetworkX快速实现(简洁易上手)
NetworkX是专门处理图结构的工具库,能直接帮我们找出所有连通的节点集合,适合快速验证需求。
步骤1:先装依赖(没装过的话)
pip install networkx
步骤2:完整代码实现
import pandas as pd import networkx as nx # 你的原始DataFrame(我补全了B列末尾的缺失值,你可以根据实际数据调整) df = pd.DataFrame({ 'A':[0, 1, 3, 4, 6, 7, 8, 8, 8, 9, 9, 9, 9, 9, 11, 12, 13, 14, 15, 15, 15, 16, 16, 16, 16, 17, 17, 17, 17, 18, 18, 18, 18, 18, 19, 19, 19, 19, 20, 20, 21, 22, 24, 25, 26, 27, 28, 29, 29], 'B':[1, 0, 4, 3, 7, 6, 112, 9, 114, 134, 135, 112, 8, 114, 14, 13, 12, 11, 16, 17, 18, 17, 15, 18, 19, 16, 18, 15, 19, 17, 16, 15, 19, 20, 20, 18, 17, 16, 19, 18, 22, 21, 25, 24, 27, 26, 29, 28, 28] }) # 1. 构建无向图:把每一行的A和B看作一条连接的边 G = nx.Graph() edges = list(zip(df['A'], df['B'])) G.add_edges_from(edges) # 2. 找出所有连通分量,给每个分量分配唯一组ID component_id = {} for group_num, component in enumerate(nx.connected_components(G)): for num in component: component_id[num] = group_num # 3. 把组ID映射回原DataFrame df['A_group'] = df['A'].map(component_id) df['B_group'] = df['B'].map(component_id) # 也可以直接给每行标记组(因为A和B属于同一组,用A或B的组ID都可以) df['group_id'] = df['A'].map(component_id) # 查看前10行结果 print(df[['A', 'B', 'group_id']].head(10))
结果说明
运行后你会看到:
- 0和1的
group_id都是0 - 8、9、112、114、134、135的
group_id都是2 - 所有互相连通的数值都会共享同一个组ID
方法二:用并查集(Union-Find)实现(高效适合大数据)
如果你的数据集特别大(百万级以上),并查集的效率会比NetworkX高很多,它是专门处理动态连通性问题的数据结构,适合生产环境的大规模数据处理。
完整代码实现
import pandas as pd # 实现并查集类(带路径压缩优化) class UnionFind: def __init__(self): self.parent = {} def find(self, x): # 查找节点x的根节点,路径压缩减少后续查找成本 if x not in self.parent: self.parent[x] = x if self.parent[x] != x: self.parent[x] = self.find(self.parent[x]) return self.parent[x] def union(self, x, y): # 合并两个节点所在的集合 x_root = self.find(x) y_root = self.find(y) if x_root != y_root: self.parent[y_root] = x_root # 你的原始DataFrame df = pd.DataFrame({ 'A':[0, 1, 3, 4, 6, 7, 8, 8, 8, 9, 9, 9, 9, 9, 11, 12, 13, 14, 15, 15, 15, 16, 16, 16, 16, 17, 17, 17, 17, 18, 18, 18, 18, 18, 19, 19, 19, 19, 20, 20, 21, 22, 24, 25, 26, 27, 28, 29, 29], 'B':[1, 0, 4, 3, 7, 6, 112, 9, 114, 134, 135, 112, 8, 114, 14, 13, 12, 11, 16, 17, 18, 17, 15, 18, 19, 16, 18, 15, 19, 17, 16, 15, 19, 20, 20, 18, 17, 16, 19, 18, 22, 21, 25, 24, 27, 26, 29, 28, 28] }) # 初始化并查集 uf = UnionFind() # 遍历所有行,合并A和B的数值 for a, b in zip(df['A'], df['B']): uf.union(a, b) # 把根节点映射成连续的组编号(让组ID更规整) root_to_group = {root: idx for idx, root in enumerate(set(uf.find(node) for node in uf.parent))} component_id = {node: root_to_group[uf.find(node)] for node in uf.parent} # 映射回DataFrame df['group_id'] = df['A'].map(component_id) # 查看前10行结果 print(df[['A', 'B', 'group_id']].head(10))
结果说明
这个方法和NetworkX的输出完全一致,但处理超大数据集时速度会快很多,适合需要高性能的场景。
内容的提问来源于stack exchange,提问作者Kdog
相关产品推荐
相关产品推荐

