You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何基于DataFrame两列数据构造关联链并标记对应分组

这个需求本质是识别有向图中的单链连通分量,可借助pandas+networkx实现,以下是完整方案:

实现代码

import pandas as pd
import networkx as nx

# 1. 构造示例DataFrame(实际使用时替换为你的数据源即可)
data = {
    "Col1": ["Type1", "Type3", "Type8", "Type3", "Type2", "Type4", "Type6", "Type9", "Type13"],
    "Col2": ["Type2", "Type4", "Type13", "Type15", "Type6", "Type9", "Type11", "Type18", "Type20"]
}
df = pd.DataFrame(data)

# 2. 统计节点出入度,筛选无分支的有效边
out_degree = df["Col1"].value_counts()
in_degree = df["Col2"].value_counts()
# 有效边规则:起点出度=1、终点入度=1,不存在分支
df["is_valid"] = df.apply(
    lambda x: out_degree[x["Col1"]] == 1 and in_degree[x["Col2"]] == 1,
    axis=1
)

# 3. 基于有效边构建图,识别连通分量分配Chain编号
G = nx.Graph()
valid_edges = df[df["is_valid"]][["Col1", "Col2"]].values.tolist()
G.add_edges_from(valid_edges)

chain_mapper = {}
for chain_id, component in enumerate(nx.connected_components(G), start=1):
    chain_name = f"Chain{chain_id}"
    for u, v in valid_edges:
        if u in component and v in component:
            chain_mapper[(u, v)] = chain_name

# 4. 映射结果到原表
df["Chain"] = df.apply(
    lambda x: chain_mapper.get((x["Col1"], x["Col2"]), ""),
    axis=1
)
df = df.drop(columns=["is_valid"])

# 输出结果
print(df)

输出结果

Col1    Col2   Chain
0   Type1   Type2  Chain1
1   Type3   Type4  Chain2
2   Type8  Type13  Chain3
3   Type3  Type15        
4   Type2   Type6  Chain1
5   Type4   Type9  Chain2
6   Type6  Type11  Chain1
7   Type9  Type18  Chain2
8  Type13  Type20  Chain3

核心逻辑说明

  • 单链上的节点不存在分支:中间节点入度、出度都为1,链首节点出度1入度0,链尾节点入度1出度0
  • 只要节点的入度/出度大于1,说明存在分支,对应的边不属于任何单链,Chain列置空
  • 所有有效边构成的无向图中,每个连通分量对应一条独立的关联链,统一分配编号即可

内容的提问来源于stack exchange,提问作者Aditya sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 01:36:03