使用Graphs.jl构建联系人网络可视化遇顶点数问题,求可行方案
解决大ID联系人网络构建与可视化问题
核心思路:ID到连续索引的映射
Graphs.jl的图顶点依赖连续整数索引,直接用百万级大ID会创建大量空顶点,完全没必要。正确做法是把所有唯一的人员ID映射成从1开始的连续小索引,这样图的顶点数就等于实际存在的人员数(最多136个,因为数据集只有68条边)。
具体实现步骤
1. 生成ID到索引的映射
先收集所有出现的person1和person2的ID,去重后给每个ID分配唯一的小索引:
using DataFrames, Graphs, SimpleWeightedGraphs # 假设你的数据在DataFrame df1里 all_ids = unique(vcat(df1.person1, df1.person2)) id_to_idx = Dict(id => i for (i, id) in enumerate(all_ids))
2. 转换边的源/目标为小索引
把原来的大ID替换成对应的索引,再构建加权图:
sources = [id_to_idx[id] for id in df1.person1] destinations = [id_to_idx[id] for id in df1.person2] weights = df1.overlaptime g = SimpleWeightedGraph(sources, destinations, weights)
3. 关联人员属性并设置顶点颜色
假设你有包含人员属性的DataFrame person_attr,列包括person_id、income_level(如"high"/"low"):
# 为每个顶点(索引)匹配对应属性 vertex_attr = [person_attr[person_attr.person_id .== all_ids[i], :income_level][1] for i in 1:nv(g)] # 映射颜色:高收入用红色,低收入用蓝色 color_map = Dict("high" => :red, "low" => :blue) vertex_colors = [color_map[attr] for attr in vertex_attr]
4. 可视化网络
推荐两种可视化方案,按需选择:
快速绘图:GraphPlot.jl
using GraphPlot gplot(g, nodefillc=vertex_colors, edgelinewidth=[w/1000 for w in weights], # 按重叠时长缩放边宽 nodelabel=all_ids, # 可选:显示原始人员ID layout=spring_layout)
定制化绘图:Makie.jl
using Makie, NetworkLayout # 计算节点布局位置 pos = spring(g) # 创建画布与轴 fig, ax = scatter(pos[:,1], pos[:,2], color=vertex_colors, markersize=15) # 绘制边 for (s, d, w) in zip(sources, destinations, weights) lines!(ax, [pos[s,1], pos[d,1]], [pos[s,2], pos[d,2]], linewidth=w/500, color=:gray) end # 添加节点标签 for (i, id) in enumerate(all_ids) text!(ax, pos[i,1], pos[i,2], text=string(id), offset=(5,5)) end display(fig)
额外提示
- 若属性数据量大,建议用
leftjoin将属性合并到索引映射的DataFrame中,避免循环查找 - 重叠时长作为边权重,可通过边宽度、透明度体现权重差异
- 如需交互式可视化,可尝试
InteractiveGraphs.jl
内容的提问来源于stack exchange,提问作者Chao
相关产品推荐
相关产品推荐

