使用R语言ggalluvial包绘制桑基图报错,请求技术帮助
问题描述
我是R语言新手,需要用ggalluvial包绘制桑基图,展示空间数据集中Cluster3(取值1-3)向Cluster6(取值1-6)的拆分关系。数据包含ID(连续值1-13336)、Cluster3、Cluster6及x/y坐标。运行代码时出现以下报错:
Error: Continuous value supplied to discrete scale In addition 1: In to_lodes_form(data = data, axes = axis_ind, discern = params$discern) : Some strata appear at multiple axes. 2: In to_lodes_form(data = data, axes = axis_ind, discern = params$discern) : Some strata appear at multiple axes. 3: In to_lodes_form(data = data, axes = axis_ind, discern = params$discern) : Some strata appear at multiple axes.
怀疑是ID为连续值导致问题,改用x/y坐标作为标识仍出现相同报错。此前尝试networkd3包但其中igraph无法正常工作,当前使用的ggalluvial代码如下:
library(ggalluvial) library(raster) library(dplyr) library(tidyverse) Clust3 <- raster(paste0("//.....3Cluster.tif")) Clus6 <- raster(paste0("//.....6Cluster.tif")) stack_sdf3 = rasterToPoints(Clust3, spatial = T) stack_df3 = as.data.frame(stack_sdf3) stack_df3$ID <- sprintf("%03d", 1:nrow(stack_df3)) stack_sdf6 = rasterToPoints(Clus6, spatial = T) stack_df6 = as.data.frame(stack_sdf6) stack_df6$ID <- sprintf("%03d", 1:nrow(stack_df6)) # Edit the column names for better understanding colnames(stack_df3) = c("Cluster_Number3","x","y","ID") colnames(stack_df6) = c("Cluster_Number6","x","y","ID") df3 <- stack_df3[, c("ID", "Cluster_Number3","x","y")] df6 <- stack_df6[, c("ID", "Cluster_Number6","x","y")] data <- df3 %>% right_join(df6, by=c("x", "y")) data2<- select(data,"ID.x", "Cluster_Number3", "Cluster_Number6") ggplot(data = data2, aes(axis1 = Cluster_Number3, axis2 = Cluster_Number6, y = ID.x)) + geom_alluvium(aes(fill = Cluster_Number3)) + geom_stratum() + geom_text(stat = "stratum", aes(label = after_stat(stratum))) + scale_x_discrete(limits = c("Main Cluster", "Diverging clusters"), expand = c(0.15, 0.05)) + scale_fill_viridis_d()+ theme_void()
解决方案
问题根源
- y轴变量错误:桑基图的
y轴需要表示每个类别组合的频数/数量,你使用的ID.x是单个样本的标识,属于离散个体值,而非聚合后的数量,导致离散刻度无法识别连续/个体值,触发Continuous value supplied to discrete scale报错。 - 聚类变量类型错误:
Cluster_Number3和Cluster_Number6是数值型变量,ggalluvial需要离散的因子类型来识别类别,否则会被误判为连续值,同时引发Some strata appear at multiple axes的警告。
修正代码
library(ggalluvial) library(raster) library(dplyr) library(tidyverse) # 读取栅格数据 Clust3 <- raster(paste0("//.....3Cluster.tif")) Clus6 <- raster(paste0("//.....6Cluster.tif")) # 转换为数据框(无需额外生成ID) stack_df3 <- as.data.frame(rasterToPoints(Clust3)) colnames(stack_df3) <- c("x", "y", "Cluster_Number3") stack_df6 <- as.data.frame(rasterToPoints(Clus6)) colnames(stack_df6) <- c("x", "y", "Cluster_Number6") # 按x/y合并数据 data <- stack_df3 %>% right_join(stack_df6, by = c("x", "y")) # 统计每个聚类组合的样本数,同时将聚类变量转为因子 data_summary <- data %>% mutate( Cluster_Number3 = factor(Cluster_Number3), Cluster_Number6 = factor(Cluster_Number6) ) %>% count(Cluster_Number3, Cluster_Number6, name = "count") # 绘制桑基图 ggplot(data = data_summary, aes(axis1 = Cluster_Number3, axis2 = Cluster_Number6, y = count)) + geom_alluvium(aes(fill = Cluster_Number3)) + geom_stratum() + geom_text(stat = "stratum", aes(label = after_stat(stratum))) + scale_x_discrete(limits = c("Main Cluster", "Diverging clusters"), expand = c(0.15, 0.05)) + scale_fill_viridis_d() + theme_void()
关键修改说明
- 去掉了不必要的ID生成,直接按x/y合并栅格数据
- 新增
data_summary步骤:统计每个(Cluster3, Cluster6)组合的样本数count,同时将聚类变量转为因子类型 - ggplot映射中用
count作为y轴变量,对应桑基图的流量大小 - 因子类型的聚类变量解决了离散刻度的报错和strata重复的警告
内容的提问来源于stack exchange,提问作者Tammy
相关产品推荐
相关产品推荐

