You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何绘制百万级元组对构成的超大规模网络图?

处理百万级边的网络图绘制问题

我有一个包含超过100万个元组对的超大边集合,示例如下:

out = {(37785, 154529),
       (37785, 349968),
       (37785, 232661),
       (185781, 72031),
       (329262, 292879),
       (380659, 183120),
       (122757, 105545),
       (160550, 6687),
       (277539, 28017),
       (345020, 166458),
       (86588, 92855),
       (94333, 37785),
       (117888, 232202),
       (258264, 156963),
       (15641, 188530),
       (96120, 399079),
       (16366, 44675),
       (321372, 37103),
       (236213, 340521) ........}

使用以下代码运行时,耗时极长且无输出结果:

import pandas as pd
import networkx as nx
import matplotlib.pyplot as plt

df = pd.DataFrame(out, columns=['from', 'to'])
G = nx.from_pandas_edgelist(df, source='from', target='to')

# 此操作将返回一个元组列表,每个元组为(node, deg)
list_degree = list(G.degree())

# 构建节点列表及对应的度数列表
nodes, degree = map(list, zip(*list_degree))

nx.draw(G, nodelist=nodes, node_size=[(v * 5) + 1 for v in degree])
plt.show()  # 绘制网络图

问题根源

直接用matplotlib绘制百万级规模的图完全不现实:

  • matplotlib是面向小数据集的可视化工具,百万级节点/边会导致渲染计算量爆炸,耗时极长
  • 节点会严重重叠成一团,无法呈现任何有效结构信息
  • 内存占用过高,甚至可能触发内存溢出

优化方案

1. 先做图的统计分析(优先)

百万级的图直接可视化没有实际意义,先通过统计提取关键信息:

import networkx as nx

# 跳过pandas,直接从边集合创建图,节省内存和时间
G = nx.Graph(out)

# 基础统计信息
print(f"节点总数: {G.number_of_nodes()}")
print(f"边总数: {G.number_of_edges()}")
print(f"平均节点度数: {sum(d for n, d in G.degree())/G.number_of_nodes():.2f}")

# 获取度数最高的10个节点
top_nodes = sorted(G.degree(), key=lambda x: x[1], reverse=True)[:10]
print("Top 10 高连通节点:", top_nodes)

# 过滤低度数节点,生成简化子图(比如保留度数≥5的节点)
filtered_nodes = [n for n, d in G.degree() if d >= 5]
filtered_G = nx.subgraph(G, filtered_nodes)
print(f"过滤后节点数: {filtered_G.number_of_nodes()}, 边数: {filtered_G.number_of_edges()}")

2. 选择合适的可视化方式

如果一定要可视化,避免用matplotlib默认绘制,改用以下方案:

方案一:抽样可视化(networkx+matplotlib)

只渲染高连通节点及其邻接边,减少渲染压力:

import matplotlib.pyplot as plt

# 选取Top 200高连通节点,以及它们的邻居节点
top_200_nodes = [n for n, d in sorted(G.degree(), key=lambda x: x[1], reverse=True)[:200]]
neighbor_nodes = set()
for node in top_200_nodes:
    neighbor_nodes.update(G.neighbors(node))

# 合并关键节点与邻居,生成抽样子图
sample_nodes = set(top_200_nodes) | neighbor_nodes
sample_G = nx.subgraph(G, sample_nodes)

# 用更高效的布局算法调整节点间距
pos = nx.spring_layout(sample_G, k=0.15, iterations=20)
# 根据度数设置节点大小
node_sizes = [d * 10 for n, d in sample_G.degree()]

# 绘制图,弱化边的颜色避免干扰
nx.draw(sample_G, pos=pos, node_size=node_sizes, alpha=0.6, edge_color="#cccccc")
plt.show()

方案二:PyVis交互式可视化

PyVis支持大图的交互式探索,生成HTML文件后可缩放、拖拽查看:

from pyvis.network import Network

# 创建网络对象,设置可视化参数
net = Network(notebook=False, height="800px", width="100%", bgcolor="#222222", font_color="white")

# 添加节点(大小基于度数)和边
for node, degree in G.degree():
    net.add_node(node, size=degree * 2, title=f"节点ID: {node}\n度数: {degree}")
for u, v in G.edges():
    net.add_edge(u, v)

# 调整物理布局参数,减少节点重叠
net.barnes_hut(gravity=-80000, central_gravity=0.3, spring_length=200, spring_strength=0.001)
# 保存为可交互的HTML文件
net.show("large_network.html")

3. 移除不必要的pandas中转

原代码用pandas将边集合转为DataFrame完全多余,直接用nx.Graph(out)创建图,能节省大量内存和转换时间。


内容的提问来源于stack exchange,提问作者ASKing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 21:20:33