You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

geom_area堆叠面积图y值超1,设ylim后出现空白线问题排查

问题

我给文章中每个句子打了标签,想生成堆叠面积图展示特定相对位置下各标签的占比:

  • 相对位置计算公式:sentence_index/total_number_of_sentence
  • 占比计算公式:位置X下,带标签A的句子总数/总句子数

我已经验证每个位置的所有标签占比总和为1,以下是loc(0.24,0.28)区间的完整数据子集:

> area_df[area_df$loc>0.24,]
    label percentage  loc
186   B1      0.195 0.25
187   C1      0.111 0.25
188   E1      0.006 0.25
189   G1      0.075 0.25
190   H1      0.008 0.25
191   M1      0.125 0.25
192   M2      0.064 0.25
193   M3      0.084 0.25
194   O1      0.070 0.25
195   O2      0.053 0.25
196   R1      0.209 0.25
197   B1      0.500 0.26
198   M2      0.250 0.26
199   M3      0.250 0.26
200   B1      0.166 0.27
201   C1      0.177 0.27
202   E1      0.015 0.27
203   G1      0.100 0.27
204   H1      0.011 0.27
205   M1      0.114 0.27
206   M2      0.048 0.27
207   M3      0.059 0.27
208   O1      0.074 0.27
209   O2      0.026 0.27
210   R1      0.210 0.27
211   B1      0.125 0.28
212   C1      0.250 0.28
213   G1      0.125 0.28
214   H1      0.125 0.28
215   M1      0.125 0.28
216   O1      0.125 0.28
217   O2      0.125 0.28

但用ggplot2的geom_area绘图时,部分位置的y值总和超过1;设置ylim(0,1)后,图中出现奇怪的空白线。我的代码如下:

# all data stored in area_df
normal_loc_uniq <- sort(unique(normal_loc))
area_df <- data.frame(matrix(ncol = 3,nrow=0))
colnames(area_df) <- c("loc","label","percentage")

# for each location, calculate the percentage
for (one_loc in normal_loc_uniq){
  subset <- data[data$normal_loc == one_loc,]
  subset_count <- as.data.frame(round(prop.table(table(subset$normal_label, useNA = "no")),5))
  names(subset_count) <- c("label","percentage")
  subset_count$loc <- as.numeric(one_loc)
  subset_count$percentage <- round(subset_count$percentage,3)

# test if there are locations with percentage not equal to 1
  if (0.98>sum(subset_count$percentage)| sum(subset_count$percentage) >1.02){
    print("error. total percentage is not 1")
  }
  area_df <- rbind(area_df,subset_count)
  }

library(ggplot2)
colors <- c("#1f77b4", "#ff7f0e", "#2ca02c", "#d62728", "#9467bd", "#8c564b", "#e377c2", "#7f7f7f", "#bcbd22", "#17becf", "#aaffc3")
ggplot(area_df, aes(x = loc, y = percentage, fill = label)) +
  geom_area(na.rm=TRUE,position="stack") + 
  scale_fill_manual(values=colors) + 
  labs(x = "Relative Location", y = "Percentage", fill = "Label") +
  theme_bw()

请排查问题原因。


问题排查与解决

核心原因1:缺失标签的0值未填充

观察你的数据子集,比如loc=0.26只有B1、M2、M3三个标签,其他标签(如C1、E1等)在该位置没有对应的行,即缺失了这些标签的占比0值。geom_area堆叠时会把缺失值当作NA处理,na.rm=TRUE会直接跳过这些缺失点,导致ggplot插值计算出现偏差,最终堆叠总和超过1,或出现空白线。

核心原因2:数据排序不规范

geom_area要求每个x值对应的所有分组(标签)数据连续且有序,如果数据未按loc和label排序,堆叠计算时会出现错位,导致总和异常。

解决方法

  1. 填充缺失标签的0值
    使用tidyr::complete函数为每个loc补充所有标签的行,缺失的percentage填充为0:

    library(tidyr)
    # 获取所有唯一标签
    all_labels <- unique(area_df$label)
    # 填充缺失值
    area_df_complete <- area_df %>%
      complete(loc, label = all_labels, fill = list(percentage = 0))
    
  2. 规范数据排序
    确保数据按loc和label排序,避免堆叠错位:

    area_df_complete <- area_df_complete %>% arrange(loc, label)
    
  3. 重新绘图
    用处理后的完整数据绘图,无需设置ylim(0,1),堆叠总和会自然保持在1:

    ggplot(area_df_complete, aes(x = loc, y = percentage, fill = label)) +
      geom_area(position = "stack") + 
      scale_fill_manual(values=colors) + 
      labs(x = "Relative Location", y = "Percentage", fill = "Label") +
      theme_bw()
    

额外验证

处理后可再次检查每个loc的总和:

area_df_complete %>%
  group_by(loc) %>%
  summarise(total = sum(percentage)) %>%
  filter(total < 0.99 | total > 1.01)

正常情况下不会返回任何行,说明每个位置的总和都正确。

内容的提问来源于stack exchange,提问作者FewKey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 07:08:08