如何将出版物作者表转换为igraph可用的合著网络边列表?
Hey there! Let's break down how to turn your publication data into a weighted edge list that works seamlessly with igraph—perfect for building that co-authorship network, even with your 7000+ entries. Here's a straightforward, efficient approach using pandas (it’ll handle the volume easily):
Step 1: Load and prep your data
First, we’ll get your data into a pandas DataFrame. If your data is in a text/csv file, use pd.read_csv; for the sample you provided, we can load it directly as a string:
import pandas as pd from itertools import combinations # Load your data (replace this with pd.read_csv("your_file.csv", sep=r'\s+') for actual files) sample_data = """PubID Author 169759 ZJ 174843 RA 174843 DJ 174843 JP 174843 GS 174843 Tv 171051 MC 171051 JR 171051 CW 171719 PB 171719 MD 171719 FO 169759 FO 173847 RA 173847 DJ""" df = pd.read_csv(pd.compat.StringIO(sample_data), sep=r'\s+')
Step 2: Generate co-authorship pairs
We’ll group authors by their PubID (since everyone on the same publication is a co-author), then create all unique, undirected pairs of authors for each group. Sorting the authors in each pair ensures we don’t get duplicate edges like RA-DJ and DJ-RA:
# Create all unique author pairs per publication edges = df.groupby('PubID')['Author'].apply( lambda authors: list(combinations(sorted(authors), 2)) ).explode() # Split the pairs into separate Source and Target columns edges = edges.apply(pd.Series).rename(columns={0: 'Source', 1: 'Target'})
Step 3: Add weights (count co-authorships)
Now we’ll count how many times each pair has collaborated—this becomes the edge weight (like RA and DJ having a weight of 2 in your sample):
weighted_edges = edges.groupby(['Source', 'Target']).size().reset_index(name='Weight')
Step 4: Use with igraph
You can either save this as a CSV file for later use, or load it directly into igraph:
# Save to a CSV (optional) weighted_edges.to_csv('coauthorship_edge_list.csv', index=False) # Load directly into igraph import igraph as ig # Create the graph with weighted edges (undirected, since co-authorship is mutual) coauthorship_graph = ig.Graph.TupleList( weighted_edges.itertuples(index=False), weights=True, directed=False )
Quick notes:
- The
sorted(authors)in the combinations step is key—it ensures we treatRA-DJandDJ-RAas the same edge, so their collaboration count is combined correctly. - Pandas handles large datasets (7000+ IDs) efficiently, so this won’t bog down your system.
- The final
weighted_edgesDataFrame has exactly what igraph needs: source, target, and weight columns.
内容的提问来源于stack exchange,提问作者otter77

