networkD3中sankeyNetwork函数如何确定节点x轴位置
R使用networkD3绘制桑基图节点排序异常问题及解决方案
问题背景
我正在查阅相关文档与教程,尝试在R中使用networkD3::sankeyNetwork()绘制桑基图。参考的Stack Overflow实现代码可以正常运行,但自行编写的代码运行后节点在x轴上排序错误,导致流向完全无法解读。
我无法确定sankeyNetwork从何处获取节点x轴位置的相关信息,以下是我未得到预期结果的实现代码:
library(tidyverse) library(networkD3) #Create the data df <- data.frame('one' = c('a', 'b', 'b', 'a'), 'two' = c('c', 'd', 'e', 'c'), 'three' = c('f', 'g', 'f', 'f')) #My code #Create the links links <- df %>% mutate(row = row_number()) %>% #Get row for grouping and pivoting pivot_longer(-row) %>% #pivot to long format group_by(row) %>% mutate(source_c = lead(value)) %>% #Get flow filter(!is.na(source_c)) %>% #Get rid of NA rename(target_c = value) %>% #Correct names group_by(target_c, source_c) %>% #Count frequencies summarize(value = n()) %>% ungroup() %>% mutate(target = as.integer(factor(target_c)), #Convert to numeric values source = as.integer(factor(source_c))) %>% mutate(source = source - 1, #zero index target = target - 1) %>% data.frame() #create the nodes nodes <- data.frame(name = factor(unique(c(links$target_c, links$source_c)))) #plot the network sankeyNetwork(Links = links, Nodes = nodes, Source = 'source', Target = 'target', Value = 'value', NodeID = 'name')
运行后错误输出:
参考的可运行代码
links <- df %>% mutate(row = row_number()) %>% # add a row id gather('col', 'source', -row) %>% # gather all columns mutate(col = match(col, names(df))) %>% # convert col names to col nums mutate(source = paste0(source, '_', col)) %>% # add col num to node names group_by(row) %>% arrange(col) %>% mutate(target = lead(source)) %>% # get target from following node in row ungroup() %>% filter(!is.na(target)) %>% # remove links from last column in original data select(source, target) %>% group_by(source, target) %>% summarise(value = n()) # aggregate and count similar links # create nodes data frame from unque nodes found in links data frame nodes <- data.frame(id = unique(c(links$source, links$target)), stringsAsFactors = FALSE) # remove column id from names nodes$name <- sub('_[0-9]*$', '', nodes$id) # set links data to the 0-based index of the nodes in the nodes data frame links$source <- match(links$source, nodes$id) - 1 links$target <- match(links$target, nodes$id) - 1 sankeyNetwork(Links = links, Nodes = nodes, Source = 'source', Target = 'target', Value = 'value', NodeID = 'name')
运行后正确输出:
疑问
我知道两份代码存在差异,但我看不出sankeyNetwork在哪里调用了代表x轴位置的列编号相关信息,全程没有看到对这类变量的引用。我希望了解符合要求的输入数据结构是什么样的,这样我就可以调整自己的数据预处理代码正常运行。
问题解答
核心原理
sankeyNetwork没有显式传入x轴位置的参数,它会自动基于有向链路的流向逻辑推断每个节点的层级(即x轴位置):只要存在链路A→B,算法就默认A的层级低于B,无入度的节点默认放在最左侧第一层,无出度的节点放在最右侧最后一层。
错误原因
你在构建节点表时直接对节点名称做了全局去重,导致不同层级的同名节点被合并为同一个节点,直接打破了流向逻辑,算法无法正确识别节点层级,最终出现排序混乱。
符合要求的输入数据结构规则
- 链路表(Links)必须是0索引的有向边集合,必须包含
source(源节点索引,整数)、target(目标节点索引,整数)、value(边的权重,数值)三个核心列 - 节点表(Nodes)的行数必须等于所有独立节点的总数量,每个节点对应唯一的0起始索引,
NodeID列对应最终渲染显示的节点名称 - 关键注意点:不同层级的同名节点必须作为独立节点存在,不能合并。你可以像参考代码一样给节点名加所属层级的后缀做区分,最后在
NodeID列中去掉后缀即可,这样就能保证链路流向严格从左到右,算法可以正确推断x轴位置。
内容的提问来源于stack exchange,提问作者MorrisseyJ
相关产品推荐
相关产品推荐

