ggplot2中fill选项导致直方图密度值异常的问题求助
Ah, I've run into this exact issue before! The root cause is how ggplot computes the ..density.. statistic when you add a fill grouping: by default, it calculates density per group instead of across the entire dataset. That's why your filled histogram has inflated density values for the b=FALSE group—its density is scaled to the size of its own subset (~84% of your data) rather than the full 10,000 observations.
Here's How to Fix It
We need to force the density calculation to use the total number of observations from the entire dataset, not just each subgroup. There are two straightforward methods:
Method 1: Adjust the Y-Aesthetic Directly in geom_histogram
Replace the default ..density.. with a custom calculation that uses the total count of all observations (sum(..count..)) instead of per-group counts:
set.seed(42) x <- rnorm(10000,0,1) df <- data.frame(x=x, b=x>1) # Correct filled density histogram ggplot(df, aes(x = x)) + geom_histogram( aes(fill = b, y = ..count../sum(..count..)/..width..), position = "stack" # Keeps bins stacked like your original filled plot )
..count../sum(..count..)gives the proportion of total observations in each bin segment- Dividing by
..width..converts that proportion to density (since density = proportion / bin width)
Method 2: Precompute Stats for Full Control
If you prefer more transparency, precompute the bin counts and densities manually using dplyr, then plot with geom_col:
library(dplyr) # Match ggplot's default bin width (30 bins total) bin_width <- diff(range(df$x)) / 30 # Calculate grouped bin stats df_binned <- df %>% # Create bins aligned with ggplot's default behavior mutate(x_bin = cut(x, breaks = seq(min(x)-bin_width/2, max(x)+bin_width/2, by = bin_width))) %>% # Count observations per bin and group group_by(x_bin, b) %>% summarise(count = n(), .groups = "drop") %>% # Compute density using total dataset size mutate( total_obs = sum(count), density = count / total_obs / bin_width, # Extract midpoint of each bin for plotting x_mid = as.numeric(substr(x_bin, 2, nchar(x_bin)-1)) %>% strsplit(",") %>% sapply(function(y) mean(as.numeric(y))) ) # Plot the precomputed data ggplot(df_binned, aes(x = x_mid, y = density, fill = b)) + geom_col(position = "stack", width = bin_width)
Why Your Original Code Failed
When you use fill=b in the aes, ggplot automatically groups the data by b. The default ..density.. statistic for grouped data is:..count.. / (n_group * ..width..)
Where n_group is the total number of observations in that subgroup (e.g., ~8413 for b=FALSE).
In contrast, the ungrouped histogram uses:..count.. / (n_total * ..width..)
Where n_total is all 10,000 observations. This mismatch is why your filled plot had inflated density values.
Verification
After applying either fix, your filled histogram will:
- Match the shape and scale of your ungrouped density histogram
- Show correct density values (peak around ~0.4, matching the standard normal distribution's maximum density)
- Keep the color-coded bins you wanted
内容的提问来源于stack exchange,提问作者dixhom

