You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言geom_sankey桑基图:自定义中间列顺序并去除NA

解决ggsankey桑基图中间列自定义排序及NA去除问题

问题分析

直接全局设置node因子级别时,因不同列(x=1/2/3)的节点类型完全不同(年份、部门、描述),会导致非目标级别的节点被过滤,进而丢失大量内容。正确的做法是仅针对中间列(x=2,对应Department)的节点单独设置排序规则,同时提前清理数据去除中间列的NA。

分步解决方案

1. 清理原始数据,去除中间列NA

先修正原始数据中Department列的字符串"NA"为真实NA,再过滤掉Department为NA的行,从源头避免中间列出现NA:

devtools::install_github("davidsjoberg/ggsankey")
library(ggsankey); library(ggplot2); library(dplyr)

# 创建数据框并清理
Years <- data.frame(Year = c(rep(2010, 5), rep(2011, 12), rep(2012, 2), rep(2013, 4), rep(2014, 5), rep(2015, 5), rep(NA, 3), rep(2022, 3), rep(NA, 2)),
                    Department = c(rep("Shoes", 4), rep("Bridal", 6), rep("Maternity", 10), rep("Winter", 3), rep("Make-Up", 6), rep("Suits", 3), rep("Vehicle", 1), rep("NA", 2), rep("Games", 4), rep(NA, 2)),
                    Description = c(rep("Place on feet", 4), rep("Wedding Dresses", 3), rep("Flowerly", 3), rep("Stretchy", 5), rep("Comfortable", 5), rep("Thick Socks", 3), rep("Foundation", 3), rep("Lipstick", 3), rep("Full Gear", 3), rep("Sedan", 1), rep("Electric", 2), rep("PC", 2), rep("Console", 2), rep(NA, 2)))

# 把字符串"NA"替换为真实NA,过滤中间列NA
Years_clean <- Years %>% 
  mutate(Department = ifelse(Department == "NA", NA, Department)) %>% 
  filter(!is.na(Department))

2. 转换为桑基图格式并自定义中间列顺序

使用make_long转换数据后,仅对**x=2(中间列)**的node和next_node设置因子级别,保证其他列的节点不受影响:

# 转换为桑基图长格式,过滤其他NA节点
df_stack <- Years_clean %>% 
  make_long(Year, Department, Description) %>% 
  filter(!is.na(node))

# 定义中间列的目标排序
middle_col_order <- c("Bridal", "Make-Up", "Maternity", "Shoes", "Suits", "Winter", "Games", "Vehicle")

# 仅针对中间列节点设置因子级别,其他节点保持原类型
df_stack <- df_stack %>%
  mutate(
    node = case_when(
      x == 2 ~ factor(node, levels = middle_col_order),
      TRUE ~ as.character(node)
    ),
    next_node = case_when(
      next_x == 2 ~ factor(next_node, levels = middle_col_order),
      TRUE ~ as.character(next_node)
    )
  )

3. 绘制桑基图

使用原绘图代码即可,此时中间列会按指定顺序排列,且无NA节点:

ggplot(df_stack, aes(x = x, 
                     next_x = next_x,
                     node = node,
                     next_node = next_node, 
                     fill = factor(node), 
                     label = node,
                     color = factor(node))) + 
  geom_sankey(flow.alpha = 0.5, node.color = 1, 
              smooth = 6, width = 0.2) + 
  geom_sankey_label(size = 3.5, color = 1, fill = "white") +
  scale_fill_viridis_d(direction = -1, option = "turbo") + 
  scale_colour_viridis_d(direction = -1, option = "turbo") +
  theme_sankey(base_size = 15) +
  theme(legend.position = "none") + 
  xlab('')

关键说明

  • 仅对中间列的节点设置因子级别,避免干扰其他列的节点(年份、描述),防止内容丢失。
  • 提前在原始数据中过滤中间列NA,比在长格式中过滤更稳妥,避免破坏桑基图的连接关系。

内容的提问来源于stack exchange,提问作者Purrsia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 13:35:08