如何用Python Pandas计算基于多列关联的唯一记录数?
解决方法
这个问题本质是求图的连通分量数量——把每一行当作图的节点,只要两行在A/B/C任意一列有相同值,就给这两个节点连一条边,最终统计有多少个独立的连通组。
方法一:使用NetworkX库实现
需要先安装依赖:pip install networkx
import pandas as pd import networkx as nx # 构造示例数据 data = { 'A': ['A1', 'A1', 'A2', 'A3'], 'B': ['B1', 'B2', 'B2', 'B3'], 'C': ['C1', 'C2', 'C3', 'C3'] } df = pd.DataFrame(data) # 初始化图并添加所有行作为节点 G = nx.Graph() G.add_nodes_from(df.index) # 遍历每一列,为同值的行添加连通边 for col in df.columns: # 按列值分组,获取每组的行索引集合 value_groups = df.groupby(col).groups for idx_group in value_groups.values(): # 组内行两两连通,用add_path高效构建连通关系 if len(idx_group) >= 2: nx.add_path(G, idx_group) # 统计连通分量的数量 connected_groups = list(nx.connected_components(G)) unique_group_count = len(connected_groups) print(f"关联组的数量:{unique_group_count}") # 输出:关联组的数量:1
方法二:并查集(Union-Find)实现(无额外依赖)
import pandas as pd class UnionFind: def __init__(self, size): self.parent = list(range(size)) def find(self, x): # 路径压缩优化 if self.parent[x] != x: self.parent[x] = self.find(self.parent[x]) return self.parent[x] def union(self, x, y): # 合并两个集合 x_root = self.find(x) y_root = self.find(y) if x_root != y_root: self.parent[y_root] = x_root # 示例数据 data = { 'A': ['A1', 'A1', 'A2', 'A3'], 'B': ['B1', 'B2', 'B2', 'B3'], 'C': ['C1', 'C2', 'C3', 'C3'] } df = pd.DataFrame(data) uf = UnionFind(len(df)) # 遍历每一列,合并同值行的集合 for col in df.columns: value_groups = df.groupby(col).groups for idx_group in value_groups.values(): if len(idx_group) > 1: # 以组内第一个元素为根,合并组内其他元素 root_idx = idx_group[0] for idx in idx_group[1:]: uf.union(root_idx, idx) # 统计不同根节点的数量,即独立关联组的数量 root_nodes = [uf.find(i) for i in range(len(df))] unique_group_count = len(set(root_nodes)) print(f"关联组的数量:{unique_group_count}") # 输出:关联组的数量:1
内容的提问来源于stack exchange,提问作者Sale
相关产品推荐
相关产品推荐

