You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas如何识别DataFrame中存在公共值的关联行并拆分数据块

实现思路

这个问题本质是图论中的连通分量识别问题:

  • 将每行视作图的一个节点
  • 若任意两行存在共有的数值,则在两个节点间连一条边
  • 最终所有互相连通的节点对应的行就属于同一个关联数据块

代码实现

版本1:借助networkx快速实现(代码简洁易读)

import pandas as pd
import networkx as nx
from itertools import combinations

# 构造示例DataFrame
df = pd.DataFrame({
    'J1': [551, 551, 2, 7, 559],
    'J2': [5, 554, 554, 6, 9],
    'J3': [552, 2, 555, 557, 560],
    'J4': [553, 5, 556, 558, 561]
})

# 1. 建立「数值→出现过的行号」映射
value_to_rows = {}
for idx, row in df.iterrows():
    for val in row.values:
        value_to_rows.setdefault(val, []).append(idx)

# 2. 构建关联图
G = nx.Graph()
G.add_nodes_from(df.index)
for rows in value_to_rows.values():
    # 同一个数值出现在多个行时,这些行两两关联
    if len(rows) >= 2:
        for u, v in combinations(rows, 2):
            G.add_edge(u, v)

# 3. 提取连通分量,拆分数据块
connected_groups = nx.connected_components(G)
data_blocks = [df.loc[list(group)] for group in connected_groups]

# 打印结果
for i, block in enumerate(data_blocks, 1):
    print(f"关联数据块{i}:")
    print(block)
    print("-"*20)

版本2:纯Python实现(无第三方依赖,大数据量性能更优)

用并查集替代networkx实现关联逻辑,其余步骤和版本1一致:

from itertools import combinations
import pandas as pd

# 并查集工具类
class UnionFind:
    def __init__(self, size):
        self.parent = list(range(size))
    def find(self, x):
        if self.parent[x] != x:
            self.parent[x] = self.find(self.parent[x])
        return self.parent[x]
    def union(self, x, y):
        xr, yr = self.find(x), self.find(y)
        if xr != yr:
            self.parent[yr] = xr

# 构造示例DataFrame
df = pd.DataFrame({
    'J1': [551, 551, 2, 7, 559],
    'J2': [5, 554, 554, 6, 9],
    'J3': [552, 2, 555, 557, 560],
    'J4': [553, 5, 556, 558, 561]
})

# 1. 建立数值到行号的映射
value_to_rows = {}
for idx, row in df.iterrows():
    for val in row.values:
        value_to_rows.setdefault(val, []).append(idx)

# 2. 用并查集合并关联行
uf = UnionFind(len(df))
for rows in value_to_rows.values():
    if len(rows) >= 2:
        for u, v in combinations(rows, 2):
            uf.union(u, v)

# 3. 按根节点分组拆分数据块
groups = {}
for idx in df.index:
    root = uf.find(idx)
    groups.setdefault(root, []).append(idx)
data_blocks = [df.loc[indices] for indices in groups.values()]

输出效果

两个版本的运行结果完全一致:

关联数据块1:
    J1   J2   J3   J4
0  551    5  552  553
1  551  554    2    5
2    2  554  555  556
--------------------
关联数据块2:
   J1  J2   J3   J4
3   7   6  557  558
--------------------
关联数据块3:
    J1  J2   J3   J4
4  559   9  560  561
--------------------

内容的提问来源于stack exchange,提问作者large rod

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 22:45:05