You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言ggalluvial包绘制桑基图报错,请求技术帮助

问题描述

我是R语言新手,需要用ggalluvial包绘制桑基图,展示空间数据集中Cluster3(取值1-3)向Cluster6(取值1-6)的拆分关系。数据包含ID(连续值1-13336)、Cluster3、Cluster6及x/y坐标。运行代码时出现以下报错:

Error: Continuous value supplied to discrete scale
In addition
1: In to_lodes_form(data = data, axes = axis_ind, discern = params$discern) :
  Some strata appear at multiple axes.
2: In to_lodes_form(data = data, axes = axis_ind, discern = params$discern) :
  Some strata appear at multiple axes.
3: In to_lodes_form(data = data, axes = axis_ind, discern = params$discern) :
  Some strata appear at multiple axes.

怀疑是ID为连续值导致问题,改用x/y坐标作为标识仍出现相同报错。此前尝试networkd3包但其中igraph无法正常工作,当前使用的ggalluvial代码如下:

library(ggalluvial)
library(raster)
library(dplyr)
library(tidyverse)

Clust3 <- raster(paste0("//.....3Cluster.tif"))
Clus6 <- raster(paste0("//.....6Cluster.tif"))

stack_sdf3 = rasterToPoints(Clust3, spatial = T)
stack_df3 = as.data.frame(stack_sdf3)
stack_df3$ID <- sprintf("%03d", 1:nrow(stack_df3))

stack_sdf6 = rasterToPoints(Clus6, spatial = T)
stack_df6 = as.data.frame(stack_sdf6)
stack_df6$ID <- sprintf("%03d", 1:nrow(stack_df6))

# Edit the column names for better understanding
colnames(stack_df3) = c("Cluster_Number3","x","y","ID")
colnames(stack_df6) = c("Cluster_Number6","x","y","ID")

df3 <- stack_df3[, c("ID", "Cluster_Number3","x","y")]
df6 <- stack_df6[, c("ID", "Cluster_Number6","x","y")]

data <- df3 %>% right_join(df6, by=c("x", "y"))
data2<- select(data,"ID.x", "Cluster_Number3", "Cluster_Number6")

ggplot(data = data2,
       aes(axis1 = Cluster_Number3, axis2 = Cluster_Number6, y = ID.x)) +
  geom_alluvium(aes(fill = Cluster_Number3)) +
  geom_stratum() +
  geom_text(stat = "stratum",
            aes(label = after_stat(stratum))) +
  scale_x_discrete(limits = c("Main Cluster", "Diverging clusters"),
                   expand = c(0.15, 0.05)) +
  scale_fill_viridis_d()+
  theme_void()
解决方案

问题根源

  1. y轴变量错误:桑基图的y轴需要表示每个类别组合的频数/数量,你使用的ID.x是单个样本的标识,属于离散个体值,而非聚合后的数量,导致离散刻度无法识别连续/个体值,触发Continuous value supplied to discrete scale报错。
  2. 聚类变量类型错误:Cluster_Number3和Cluster_Number6是数值型变量,ggalluvial需要离散的因子类型来识别类别,否则会被误判为连续值,同时引发Some strata appear at multiple axes的警告。

修正代码

library(ggalluvial)
library(raster)
library(dplyr)
library(tidyverse)

# 读取栅格数据
Clust3 <- raster(paste0("//.....3Cluster.tif"))
Clus6 <- raster(paste0("//.....6Cluster.tif"))

# 转换为数据框(无需额外生成ID)
stack_df3 <- as.data.frame(rasterToPoints(Clust3))
colnames(stack_df3) <- c("x", "y", "Cluster_Number3")

stack_df6 <- as.data.frame(rasterToPoints(Clus6))
colnames(stack_df6) <- c("x", "y", "Cluster_Number6")

# 按x/y合并数据
data <- stack_df3 %>% right_join(stack_df6, by = c("x", "y"))

# 统计每个聚类组合的样本数,同时将聚类变量转为因子
data_summary <- data %>%
  mutate(
    Cluster_Number3 = factor(Cluster_Number3),
    Cluster_Number6 = factor(Cluster_Number6)
  ) %>%
  count(Cluster_Number3, Cluster_Number6, name = "count")

# 绘制桑基图
ggplot(data = data_summary,
       aes(axis1 = Cluster_Number3, axis2 = Cluster_Number6, y = count)) +
  geom_alluvium(aes(fill = Cluster_Number3)) +
  geom_stratum() +
  geom_text(stat = "stratum", aes(label = after_stat(stratum))) +
  scale_x_discrete(limits = c("Main Cluster", "Diverging clusters"),
                   expand = c(0.15, 0.05)) +
  scale_fill_viridis_d() +
  theme_void()

关键修改说明

  • 去掉了不必要的ID生成,直接按x/y合并栅格数据
  • 新增data_summary步骤:统计每个(Cluster3, Cluster6)组合的样本数count,同时将聚类变量转为因子类型
  • ggplot映射中用count作为y轴变量,对应桑基图的流量大小
  • 因子类型的聚类变量解决了离散刻度的报错和strata重复的警告

内容的提问来源于stack exchange,提问作者Tammy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 07:37:20