You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用ggplot2的stat_density_2d仅绘制分类变量的高密度区域

Hey there! I've run into this exact issue before, so let's break down how to fix it and get your density plots looking clean and readable.

First, let's recap your setup to make sure we're on the same page:

library(ggplot2)

# Your sample data
plot_data <- data.frame(
  X = c(rnorm(300, 3, 2.5), rnorm(150, 7, 2)),
  Y = c(rnorm(300, 6, 2.5), rnorm(150, 2, 2)),
  Label = c(rep('A', 300), rep('B', 150))
)

# The overlapping density plot you're starting with
ggplot(plot_data, aes(X, Y, color = Label)) +
  geom_point(alpha = 0.3) +
  stat_density2d(geom = "polygon", aes(fill = Label), alpha = 0.3)

The problem here is that default density plots include all probability levels, leading to messy overlap. Let's fix this by only keeping regions where the density is above your 0.03 threshold.

Solution 1: Filter High-Density Regions Directly with after_stat()

This is the simplest and most straightforward approach. We can use ggplot2's after_stat() function to filter out any parts of the density plot where the calculated density (..density..) is less than 0.03, right in the stat_density2d layer:

ggplot(plot_data, aes(X, Y, color = Label)) +
  geom_point(alpha = 0.3) +
  stat_density2d(
    geom = "polygon",
    aes(fill = Label),
    alpha = 0.3,
    # Only keep regions where density exceeds 0.03
    filter = after_stat(density > 0.03)
  )

This skips rendering all low-density areas entirely, no manual tweaking of ..levels.. required.

Solution 2: Manually Set Levels Based on Density Threshold

If you prefer to work with the levels parameter (maybe for more control over the number of contour bands), you can pre-calculate the probability levels that correspond to your 0.03 density threshold for each group. This is a bit more involved but works well if you need multiple high-density bands:

library(dplyr)
library(MASS)
library(purrr)

# Calculate density surfaces for each group
density_groups <- plot_data %>%
  group_split(Label) %>%
  map(~ kde2d(.$X, .$Y, n = 100))

# For each group, find the probability level where density hits 0.03, then generate upper bands
level_thresholds <- map(density_groups, function(kde) {
  dens_values <- as.vector(kde$z)
  # Get the cumulative probability at 0.03 density
  threshold_prob <- ecdf(dens_values)(0.03)
  # Create 5 bands from the threshold up to the highest density
  seq(threshold_prob, 1, length.out = 5)
})

# Plot with custom levels per group
ggplot(plot_data, aes(X, Y, color = Label)) +
  geom_point(alpha = 0.3) +
  stat_density2d(
    geom = "polygon",
    aes(fill = Label),
    alpha = 0.3,
    levels = level_thresholds
  )

Why scale_alpha_continuous Didn't Work

You mentioned trying scale_alpha_continuous with range or limits and it didn't help—this makes sense! That parameter only adjusts how transparency is mapped to density values, not which regions are drawn. It would make low-density areas more transparent, but they'd still be there cluttering up your plot. The filter approach above actually removes those regions entirely, which is what you need.

内容的提问来源于stack exchange,提问作者Jonas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:23:32