You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将出版物作者表转换为igraph可用的合著网络边列表?

Hey there! Let's break down how to turn your publication data into a weighted edge list that works seamlessly with igraph—perfect for building that co-authorship network, even with your 7000+ entries. Here's a straightforward, efficient approach using pandas (it’ll handle the volume easily):

Step 1: Load and prep your data

First, we’ll get your data into a pandas DataFrame. If your data is in a text/csv file, use pd.read_csv; for the sample you provided, we can load it directly as a string:

import pandas as pd
from itertools import combinations

# Load your data (replace this with pd.read_csv("your_file.csv", sep=r'\s+') for actual files)
sample_data = """PubID Author
169759 ZJ
174843 RA
174843 DJ
174843 JP
174843 GS
174843 Tv
171051 MC
171051 JR
171051 CW
171719 PB
171719 MD
171719 FO
169759 FO
173847 RA
173847 DJ"""

df = pd.read_csv(pd.compat.StringIO(sample_data), sep=r'\s+')

Step 2: Generate co-authorship pairs

We’ll group authors by their PubID (since everyone on the same publication is a co-author), then create all unique, undirected pairs of authors for each group. Sorting the authors in each pair ensures we don’t get duplicate edges like RA-DJ and DJ-RA:

# Create all unique author pairs per publication
edges = df.groupby('PubID')['Author'].apply(
    lambda authors: list(combinations(sorted(authors), 2))
).explode()

# Split the pairs into separate Source and Target columns
edges = edges.apply(pd.Series).rename(columns={0: 'Source', 1: 'Target'})

Step 3: Add weights (count co-authorships)

Now we’ll count how many times each pair has collaborated—this becomes the edge weight (like RA and DJ having a weight of 2 in your sample):

weighted_edges = edges.groupby(['Source', 'Target']).size().reset_index(name='Weight')

Step 4: Use with igraph

You can either save this as a CSV file for later use, or load it directly into igraph:

# Save to a CSV (optional)
weighted_edges.to_csv('coauthorship_edge_list.csv', index=False)

# Load directly into igraph
import igraph as ig

# Create the graph with weighted edges (undirected, since co-authorship is mutual)
coauthorship_graph = ig.Graph.TupleList(
    weighted_edges.itertuples(index=False),
    weights=True,
    directed=False
)

Quick notes:

  • The sorted(authors) in the combinations step is key—it ensures we treat RA-DJ and DJ-RA as the same edge, so their collaboration count is combined correctly.
  • Pandas handles large datasets (7000+ IDs) efficiently, so this won’t bog down your system.
  • The final weighted_edges DataFrame has exactly what igraph needs: source, target, and weight columns.

内容的提问来源于stack exchange,提问作者otter77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:42:38