You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将表格按直接/间接关联元素分组?

基于Pandas的连通分量分组方案

这里给你两种优雅的实现方式,替代单纯的for循环,完美满足你的分组需求:

方案一:结合NetworkX(简洁高效)

NetworkX是Python专门处理图结构的库,和Pandas兼容性极佳,能快速找出所有直接/间接关联的节点分组:

首先安装依赖(如果未安装):
pip install networkx

代码实现:

import pandas as pd
import networkx as nx

# 构造你的表格数据
data = [
    ["a", "x"],
    ["b", "x"],
    ["c", "y"],
    ["c", "z"],
    ["d", "x"],
    ["x", "a"],
    ["x", "x"],
    ["y", "z"]
]
df = pd.DataFrame(data, columns=["A", "B"])

# 构建无向图,自动处理双向边,过滤自环(不影响结果但更干净)
G = nx.from_pandas_edgelist(df, source="A", target="B", create_using=nx.Graph())
G.remove_edges_from(nx.selfloop_edges(G))

# 获取所有连通分量,转成集合列表
connected_groups = [set(component) for component in nx.connected_components(G)]

print(connected_groups)
# 输出:[{'a', 'b', 'd', 'x'}, {'c', 'y', 'z'}]

说明:nx.from_pandas_edgelist直接将DataFrame的A、B列转为图的边,无向图会自动合并重复的双向边;nx.connected_components直接返回所有连通的节点集合,完全匹配你的分组要求。

方案二:并查集(Union-Find)算法(无需额外库)

如果不想引入第三方库,可以用经典的并查集算法配合Pandas实现,逻辑清晰且高效:

import pandas as pd

# 实现并查集类
class UnionFind:
    def __init__(self):
        self.parent = {}
    
    def find(self, x):
        if self.parent[x] != x:
            self.parent[x] = self.find(self.parent[x])  # 路径压缩优化
        return self.parent[x]
    
    def union(self, x, y):
        # 初始化节点
        if x not in self.parent:
            self.parent[x] = x
        if y not in self.parent:
            self.parent[y] = y
        # 合并两个节点的连通分量
        root_x = self.find(x)
        root_y = self.find(y)
        if root_x != root_y:
            self.parent[root_y] = root_x

# 构造表格数据
data = [
    ["a", "x"],
    ["b", "x"],
    ["c", "y"],
    ["c", "z"],
    ["d", "x"],
    ["x", "a"],
    ["x", "x"],
    ["y", "z"]
]
df = pd.DataFrame(data, columns=["A", "B"])

# 初始化并查集
uf = UnionFind()

# 遍历每一行,合并A和B的元素
for _, row in df.iterrows():
    a, b = row["A"], row["B"]
    if a != b:  # 跳过自环
        uf.union(a, b)

# 按根节点分组,生成连通分量
groups = {}
for node in uf.parent:
    root = uf.find(node)
    groups.setdefault(root, []).append(node)

# 转成集合列表
connected_groups = [set(group) for group in groups.values()]

print(connected_groups)
# 输出:[{'a', 'b', 'd', 'x'}, {'c', 'y', 'z'}]

说明:并查集是处理连通性问题的经典算法,通过find找根节点、union合并分量,遍历DataFrame完成所有节点的关联后,按根节点分组即可得到目标结果,全程无需依赖第三方库。

内容的提问来源于stack exchange,提问作者Mike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 06:50:16