如何从Pandas邻接矩阵DataFrame创建正确的NetworkX有向图?
I’ve dealt with this exact frustration before when working with directed graphs and pandas adjacency matrices—let’s break down what’s going wrong and how to fix it properly.
Your two approaches are hitting problems because networkx.from_pandas_adjacency is optimized for undirected graphs by default, which clashes with your need to preserve directed edge weights:
- Option A: When you first create an undirected graph, NetworkX collapses each pair of nodes (like A-B) into a single undirected edge, keeping only one weight value (often the last one processed or aggregated). Converting this to a
DiGraphdoesn’t restore the original directed edges—it just creates directed edges based on the collapsed undirected ones, so you lose one of the original directed edges entirely. - Option B: Even with
create_using=networkx.DiGraph(), the underlying logic offrom_pandas_adjacencystill carries over undirected graph assumptions. In your case, this leads to incorrect weight propagation, where one edge’s weight gets applied to its reverse direction too.
The most reliable way to preserve all your directed edge weights is to avoid relying on from_pandas_adjacency for directed graphs. Instead, convert your adjacency matrix to an edge list first, or manually add edges to the graph. Here are two straightforward solutions:
Approach 1: Convert to Edge List with Pandas stack()
This uses pandas' stack() method to reshape your adjacency matrix into a long-format edge list, which NetworkX can process correctly for directed graphs:
import pandas as pd import networkx as nx # Your original adjacency DataFrame df = pd.DataFrame( data=[[0, 0.5, 0.5, 0], [1, 0, 0, 0], [0.8, 0, 0, 0.2], [0, 0, 1, 0]], index=['A', 'B', 'C', 'D'], columns=['A', 'B', 'C', 'D'] ) # Reshape adjacency matrix to edge list edgelist = df.stack().reset_index() edgelist.columns = ['source', 'target', 'weight'] # Optional: Filter out edges with 0 weight (remove if you want to keep them) edgelist = edgelist[edgelist['weight'] != 0] # Build the directed graph G = nx.from_pandas_edgelist( edgelist, source='source', target='target', edge_attr='weight', create_using=nx.DiGraph() ) # Verify the edges print(list(filter(lambda x: x[0] == 'A' and x[1] == 'B', G.edges.data()))) print(list(filter(lambda x: x[0] == 'B' and x[1] == 'A', G.edges.data())))
Expected Output:
[('A', 'B', {'weight': 0.5})] [('B', 'A', {'weight': 1})]
Approach 2: Manually Iterate Over the Matrix
If you prefer full control, loop through each cell in your DataFrame and add edges directly to the directed graph:
import pandas as pd import networkx as nx df = pd.DataFrame( data=[[0, 0.5, 0.5, 0], [1, 0, 0, 0], [0.8, 0, 0, 0.2], [0, 0, 1, 0]], index=['A', 'B', 'C', 'D'], columns=['A', 'B', 'C', 'D'] ) # Initialize empty directed graph G = nx.DiGraph() # Add all nodes from the DataFrame index/columns G.add_nodes_from(df.index) # Iterate over each source-target pair and add edges with their weights for source in df.index: for target in df.columns: weight = df.loc[source, target] if weight != 0: # Skip zero-weight edges (optional) G.add_edge(source, target, weight=weight) # Verify the edges print(list(filter(lambda x: x[0] == 'A' and x[1] == 'B', G.edges.data()))) print(list(filter(lambda x: x[0] == 'B' and x[1] == 'A', G.edges.data())))
This will produce the same correct output as the first approach.
Both methods treat each directed edge (source → target) as a separate, unique entity, preserving the exact weight from your original adjacency matrix. Unlike from_pandas_adjacency, which has built-in assumptions for undirected graphs, these approaches don’t collapse or overwrite edge weights for reverse directions.
内容的提问来源于stack exchange,提问作者Melsauce

