You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

百分比堆叠条形图未占满100%的技术求助

解决样本级100%堆叠条形图绘制问题

问题背景

需要绘制堆叠条形图,展示每个样本中各污染来源对总ASV计数的百分比贡献,要求每个样本的条形占满100%(例如3H2C样本应全部为shipcontrolsea对应的蓝色)。但当前绘图结果显示的是基于全数据集总计数的相对丰度,无法达到预期效果。

数据与当前计算逻辑

示例数据集trytry包含以下字段:

  • name:样本名称
  • Genus:ASV属
  • cont:ASV来源对照
  • contmeanabund:经decostand(x, method="total")计算的ASV相对丰度
  • perc2:当前按全数据集总和计算的百分比,计算代码如下:
sum <- sum(as.numeric(trytry$contmeanabund))
trytry$perc2 <- (as.numeric(trytry$contmeanabund) / sum) * 100

当前绘图代码

ugh <- ggplot(trytry, aes(x=name,y=perc2,fill=cont)) +
      geom_bar(position="stack", stat="identity") +
      theme_bw() +
      theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust=1)) 
print(ugh)

解决方案

问题核心是百分比计算的分组维度错误:当前perc2基于全数据集总和计算,而我们需要按**每个样本(name)**分组,计算该样本内各污染来源的占比。

步骤1:重新计算样本内百分比

使用dplyr按样本分组,计算每个样本内各来源的占比:

library(dplyr)

# 取消数据集原有分组(若存在)
trytry <- trytry %>% ungroup()

# 按样本分组,计算样本内各污染来源的百分比
trytry <- trytry %>%
  group_by(name) %>%
  mutate(sample_perc = (contmeanabund / sum(contmeanabund)) * 100) %>%
  ungroup()

步骤2:绘制100%堆叠条形图

用新计算的sample_perc替换原perc2绘图:

library(ggplot2)

ugh <- ggplot(trytry, aes(x=name, y=sample_perc, fill=cont)) +
  geom_bar(position="stack", stat="identity") +
  theme_bw() +
  theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust=1)) +
  labs(y="百分比(%)") # 可选:添加清晰的y轴标签
print(ugh)

补充优化:去重聚合

原数据集中存在重复行(例如3H1C样本有多个相同的SAR116_clade条目),建议先按name和cont分组聚合,避免重复计算:

trytry_clean <- trytry %>%
  ungroup() %>%
  group_by(name, cont) %>%
  summarise(total_contmeanabund = sum(contmeanabund), .groups = "drop") %>%
  group_by(name) %>%
  mutate(sample_perc = (total_contmeanabund / sum(total_contmeanabund)) * 100)

使用trytry_clean绘图,结果会更准确。

内容的提问来源于stack exchange,提问作者Geomicro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 16:42:46