You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ggplot2中fill选项导致直方图密度值异常的问题求助

Fixing Density Calculation for Grouped Histograms in ggplot2

Ah, I've run into this exact issue before! The root cause is how ggplot computes the ..density.. statistic when you add a fill grouping: by default, it calculates density per group instead of across the entire dataset. That's why your filled histogram has inflated density values for the b=FALSE group—its density is scaled to the size of its own subset (~84% of your data) rather than the full 10,000 observations.

Here's How to Fix It

We need to force the density calculation to use the total number of observations from the entire dataset, not just each subgroup. There are two straightforward methods:


Method 1: Adjust the Y-Aesthetic Directly in geom_histogram

Replace the default ..density.. with a custom calculation that uses the total count of all observations (sum(..count..)) instead of per-group counts:

set.seed(42)
x <- rnorm(10000,0,1)
df <- data.frame(x=x, b=x>1)

# Correct filled density histogram
ggplot(df, aes(x = x)) +
  geom_histogram(
    aes(fill = b, y = ..count../sum(..count..)/..width..),
    position = "stack"  # Keeps bins stacked like your original filled plot
  )
  • ..count../sum(..count..) gives the proportion of total observations in each bin segment
  • Dividing by ..width.. converts that proportion to density (since density = proportion / bin width)

Method 2: Precompute Stats for Full Control

If you prefer more transparency, precompute the bin counts and densities manually using dplyr, then plot with geom_col:

library(dplyr)

# Match ggplot's default bin width (30 bins total)
bin_width <- diff(range(df$x)) / 30

# Calculate grouped bin stats
df_binned <- df %>%
  # Create bins aligned with ggplot's default behavior
  mutate(x_bin = cut(x, breaks = seq(min(x)-bin_width/2, max(x)+bin_width/2, by = bin_width))) %>%
  # Count observations per bin and group
  group_by(x_bin, b) %>%
  summarise(count = n(), .groups = "drop") %>%
  # Compute density using total dataset size
  mutate(
    total_obs = sum(count),
    density = count / total_obs / bin_width,
    # Extract midpoint of each bin for plotting
    x_mid = as.numeric(substr(x_bin, 2, nchar(x_bin)-1)) %>% 
      strsplit(",") %>% 
      sapply(function(y) mean(as.numeric(y)))
  )

# Plot the precomputed data
ggplot(df_binned, aes(x = x_mid, y = density, fill = b)) +
  geom_col(position = "stack", width = bin_width)

Why Your Original Code Failed

When you use fill=b in the aes, ggplot automatically groups the data by b. The default ..density.. statistic for grouped data is:
..count.. / (n_group * ..width..)
Where n_group is the total number of observations in that subgroup (e.g., ~8413 for b=FALSE).

In contrast, the ungrouped histogram uses:
..count.. / (n_total * ..width..)
Where n_total is all 10,000 observations. This mismatch is why your filled plot had inflated density values.

Verification

After applying either fix, your filled histogram will:

  • Match the shape and scale of your ungrouped density histogram
  • Show correct density values (peak around ~0.4, matching the standard normal distribution's maximum density)
  • Keep the color-coded bins you wanted

内容的提问来源于stack exchange,提问作者dixhom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:44:52