R语言geom_sankey桑基图:节点对齐、空方块及尺寸适配问题求助
解决方案
问题分析与分步解决
- 节点按年份排序:原数据中年份存在数值/字符混合类型,导致节点排序不符合时间逻辑,需统一类型并指定排序规则。
- 移除空方块:原数据中的NA值被
make_long转换为无效节点,需过滤这类流转记录。 - 统一节点尺寸:第一列节点偏小源于有效数据占比低,过滤NA后节点高度将匹配实际流量,同时通过参数保证宽度一致。
修改后的完整代码
devtools::install_github("davidsjoberg/ggsankey") library(ggsankey); library(ggplot2); library(dplyr) # 统一年份为数值型,避免类型混合 Years <- data.frame( Earlier = c(rep(2012, 2), 2013, 2014, rep(2015, 2), rep(2018, 2), rep(2022, 2), rep(NA, 31)), Latest = c(rep(2023, 4), rep(2022, 6), rep(2021, 10), rep(2020, 3), rep(2019, 6), rep(2018, 3), rep(2017, 3), rep(2013, 4), rep(NA, 2)), Current = c(rep(2023, 10), rep(2022, 12), rep(2021, 11), rep(2020, 1), rep(NA, 7)) ) # 洗牌并转换为长格式,过滤NA相关流转 set.seed(123) Years_shuffled <- Years[sample(1:nrow(Years)), ] df_stack <- Years_shuffled %>% make_long(Earlier, Latest, Current) %>% filter(!is.na(node) & !is.na(next_node)) %>% mutate( node_num = as.numeric(node), next_node_num = as.numeric(next_node) ) # 绘图:指定节点时间排序,统一参数 ggplot(df_stack, aes( x = x, next_x = next_x, node = factor(node, levels = sort(unique(df_stack$node_num))), next_node = factor(next_node, levels = sort(unique(df_stack$node_num))), fill = factor(node_num), label = node, color = factor(node_num) )) + geom_sankey(flow.alpha = 0.5, node.color = 1, smooth = 6, width = 0.2) + geom_sankey_label(size = 3.5, color = 1, fill = "white") + scale_fill_viridis_d(direction = -1, option = "turbo") + scale_colour_viridis_d(direction = -1, option = "turbo") + theme_sankey(base_size = 15) + theme(legend.position = "none") + xlab('') + scale_x_discrete(labels = c("Earlier", "Latest", "Current"))
关键修改说明
- 统一数据类型:将所有年份改为数值型,避免字符排序逻辑干扰时间顺序。
- 过滤无效记录:用
filter移除含NA的节点流转,彻底消除空方块。 - 指定节点排序:把
node转为按年份升序排列的因子,确保节点从上到下符合时间顺序(需降序则在sort中加decreasing = TRUE)。 - 统一节点尺寸:过滤NA后第一列节点高度匹配实际流量,
width = 0.2参数保证所有节点宽度一致。
内容的提问来源于stack exchange,提问作者Purrsia
相关产品推荐
相关产品推荐

