如何通过匹配ID编号创建igraph对象并生成边列表
解决方案
你需要生成的是参与者合作网络(同一项目参与人之间存在连接),核心逻辑是对每个项目下的所有参与者ID做两两组合生成边,具体实现代码如下:
前置说明
你的原始数据属于典型的二模网络结构,包含参与者、项目两类节点,你可以按需选择生成一模参与者合作网络或者直接使用二模网络做分析。
R 语言实现(适配igraph包)
1. 依赖加载与数据预处理
# 加载依赖包 library(tidyverse) library(igraph) # 读取你的原始数据(替换为你的实际数据读取代码) df <- read.csv("你的数据文件路径.csv") # 去重:删除同一参与者在同一项目的重复记录 df_unique <- distinct(df, Participant_ID, Project_Number)
2. 生成带权重的边列表并构建igraph对象
# 按项目分组,对每组内的参与者ID两两配对生成边 edge_list <- df_unique %>% group_by(Project_Number) %>% # 过滤只有1个参与者的项目,无合作关系无需生成边 filter(n() >= 2) %>% summarise(edges = list(combn(sort(Participant_ID), 2, simplify = FALSE))) %>% unnest(edges) %>% mutate( from = map_int(edges, ~ .x[1]), to = map_int(edges, ~ .x[2]) ) %>% ungroup() %>% select(from, to) # 统计边权重:即两个参与者共同参与的项目数量 weighted_edge <- edge_list %>% count(from, to, name = "co_project_count") # 构建无向igraph对象 g <- graph_from_data_frame(weighted_edge, directed = FALSE, vertices = NULL)
Python 实现(适配python-igraph包)
1. 依赖加载与数据预处理
import pandas as pd import itertools import igraph as ig # 读取原始数据(替换为你的实际数据读取代码) df = pd.read_csv("你的数据文件路径.csv") # 去重:删除同一参与者在同一项目的重复记录 df_unique = df.drop_duplicates(subset=["Participant_ID", "Project_Number"])
2. 生成带权重的边列表并构建igraph对象
edge_raw = [] # 按项目分组遍历 for proj_id, group in df_unique.groupby("Project_Number"): participants = group["Participant_ID"].sort_values().tolist() # 过滤仅1人参与的项目 if len(participants) < 2: continue # 生成两两配对的边 for pair in itertools.combinations(participants, 2): edge_raw.append(pair) # 统计边权重 edge_df = pd.DataFrame(edge_raw, columns=["from", "to"]).value_counts().reset_index(name="co_project_count") # 构建无向igraph对象 g = ig.Graph.TupleList( edge_df[["from", "to", "co_project_count"]].itertuples(index=False), directed=False, edge_attrs=["co_project_count"] )
可选方案:直接使用二模网络内置投影功能
如果你不需要自定义边生成逻辑,也可以直接构建二模网络后调用igraph内置的投影函数得到合作网络,代码更简洁:
R版本
# 构建二模网络,type属性标记节点是参与者还是项目 df_bimodal <- df_unique %>% rename(from = Participant_ID, to = Project_Number) g_bimodal <- graph_from_data_frame(df_bimodal, directed = FALSE) V(g_bimodal)$type <- bipartite_mapping(g_bimodal)$type # 投影得到参与者合作网络 g_coop <- bipartite_projection(g_bimodal, which = "false")
Python版本
# 构建二模网络 tuples = [tuple(x) for x in df_unique[["Participant_ID", "Project_Number"]].values] g_bimodal = ig.Graph.TupleList(tuples, directed = False) # 标记节点类型 g_bimodal.vs["type"] = [v["name"] in df_unique["Participant_ID"].unique() for v in g_bimodal.vs] # 投影得到参与者合作网络 g_coop = g_bimodal.bipartite_projection(which=True)
注意事项
- 你的1500行数据量极小,上述代码运行无任何性能压力,可适配任意人数的项目
- 如果需要保留项目作为节点做二模网络分析,无需做投影,直接使用第一步生成的二模网络即可
内容的提问来源于stack exchange,提问作者RWJ96_x
相关产品推荐
相关产品推荐

