如何在R中用ggplot2生成符合要求的Sankey/Alluvial图?
使用ggplot2绘制Sankey/Alluvial图的问题解决指南
问题概述
我尝试用R的ggplot2包,基于含x、node、next_x、next_node列的CSV数据绘制Sankey/Alluvial图以展示节点流向,要求排除next_x为NA的流。但现有代码运行时出现dplyr::mutate()的NA转换警告,生成的图仅显示节点,流向连接要么缺失要么错误,调整后仍有问题,推测是行顺序导致的。另外我还有包含物种、基因、起止位置和方向的原始基因邻居数据集,需要解决以下核心问题:
- 基于
node和next_node创建正确的流向图 - 排除
next_x为NA的流 - 消除
dplyr::mutate()的NA警告 - 解决绘图不完整、流向连接错误的问题
解决方案
1. 数据预处理(根治NA警告+过滤无效流)
先彻底过滤无效数据,统一数据类型,从根源消除NA警告:
library(tidyverse) library(ggalluvial) # 适配ggplot2的Alluvial专用包 # 读取并清洗数据 df <- read.csv("your_data.csv") %>% # 直接移除next_x为NA的行,避免后续处理产生NA filter(!is.na(next_x)) %>% # 将node和next_node转为因子,统一类型,消除mutate时的类型转换警告 mutate(across(c(node, next_node), as.factor)) %>% # 按node和next_node排序,解决行顺序导致的流向错乱 arrange(node, next_node)
2. 转换为Alluvial图标准格式
Alluvial图需要长格式数据,将x和next_x作为两个阶段,匹配对应的节点并统计流向频次:
alluvial_data <- df %>% # 为每行添加计数(若原始数据无频次列) mutate(count = 1) %>% # 重塑数据:拆分阶段与对应节点 pivot_longer( cols = c(x, next_x), names_to = "stage_type", values_to = "stage_position" ) %>% # 匹配每个阶段对应的节点类别 mutate( node_category = case_when( stage_type == "x" ~ as.character(node), stage_type == "next_x" ~ as.character(next_node) ), stage_position = as.numeric(stage_position) # 确保阶段位置为数值型 ) %>% # 按阶段、节点统计总流向数 group_by(stage_position, node_category) %>% summarise(total_flow = sum(count), .groups = "drop")
3. 绘制正确的Alluvial流向图
用ggalluvial包绘制,它会自动处理流向连接逻辑,避免手动处理的错误:
ggplot(alluvial_data, aes( x = stage_position, stratum = node_category, alluvium = total_flow, y = total_flow, fill = node_category )) + # 绘制流向带 geom_alluvium(width = 1/10, alpha = 0.7) + # 绘制节点区块 geom_stratum(width = 1/10, color = "black") + # 添加节点标签 geom_text(stat = "stratum", aes(label = node_category), size = 3.5) + # 图表美化 labs( x = "阶段位置", y = "流向数量", title = "基因-物种流向可视化", fill = "节点类别" ) + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
4. 额外排查点
如果流向仍有错误,检查:
- 原始数据中是否存在同一
node对应多个next_node但未统计频次的情况,确保total_flow统计正确 - 确认
stage_position是连续/有序的数值,避免阶段顺序混乱
内容的提问来源于stack exchange,提问作者Rohan Nath
相关产品推荐
相关产品推荐

