You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自动设置散点图分割线截距,将散点划分为Top K层级

自动计算层级边界平均值作为ggplot截距

现有数据集df,其中price_index和quantity_index按层级划分,我用ggplot绘制散点图时手动设置了geom_vline和geom_hline的截距,现在需要自动计算相邻层级边界值的平均值作为截距——比如价格维度取avg(70.50442, 67.11609)和avg(74.41048, 73.24853)作为竖线截距,数量维度取avg(79.71225, 75.43498)和avg(90.95081, 89.82692)作为横线截距。

数据集预览

products price_index quantity_index   price_rank quantity_rank
1         a    95.00000       95.00000   high_price   high_volume
2         b    80.69012       94.53585   high_price   high_volume
3         c    74.41048       90.95081   high_price   high_volume
4         d    73.24853       89.82692 medium_price medium_volume
5         e    70.50442       79.71225 medium_price medium_volume
6         f    67.11609       75.43498    low_price    low_volume
7         g    64.14685       58.26419    low_price    low_volume
8         h    56.76375       56.16531    low_price    low_volume
9         i    55.76472       56.02838    low_price    low_volume
10        j    55.70475       50.24873    low_price    low_volume

当前手动设置截距的绘图代码

ggplot(df, aes(x = price_index, y = quantity_index, label = products)) +
   coord_fixed() +
   geom_point(colour = 'blue', size = 3, alpha=0.9) +
   scale_x_continuous(expand = c(0, 0), limits = c(50, 100),
                      breaks = seq(50, 100, 5)
   ) +
   scale_y_continuous(expand = c(0, 0), limits = c(50, 100),
                      breaks = seq(50, 100, 5)
   ) +
   geom_vline(xintercept = c(68, 74), color='black',
              linetype='dotted', size=1) +
   geom_hline(yintercept = c(77, 90), color='black',
              linetype='dotted', size=1)

数据集结构

df <- structure(list(products = c("a", "b", "c", "d", "e", "f", "g", 
"h", "i", "j"), price_index = c(95, 80.69011538, 74.41047705, 
73.24853055, 70.5044217, 67.1160916, 64.14685495, 56.76375355, 
55.76472446, 55.70475052), quantity_index = c(95, 94.53585227, 
90.95080999, 89.82692004, 79.71224701, 75.43498354, 58.26419203, 
56.16530529, 56.02838119, 50.24873055), price_rank = c("high_price", 
"high_price", "high_price", "medium_price", "medium_price", "low_price", 
"low_price", "low_price", "low_price", "low_price"), quantity_rank = c("high_volume", 
"high_volume", "high_volume", "medium_volume", "medium_volume", 
"low_volume", "low_volume", "low_volume", "low_volume", "low_volume"
)), class = "data.frame", row.names = c(NA, -10L))

解决方法

步骤1:自动计算截距

通过分组提取各层级的边界值,再计算相邻边界的平均值:

library(dplyr)

# 计算价格维度截距:取各price_rank组的最小price_index,再算相邻组边界的平均值
price_cuts <- df %>%
  group_by(price_rank) %>%
  summarise(bound = min(price_index)) %>%
  arrange(desc(bound)) %>% # 按边界从高到低排序,确保层级对应正确
  pull(bound) %>%
  { (.[-1] + .[-length(.)]) / 2 } # 相邻值相加取平均

# 计算数量维度截距:逻辑同上
quantity_cuts <- df %>%
  group_by(quantity_rank) %>%
  summarise(bound = min(quantity_index)) %>%
  arrange(desc(bound)) %>%
  pull(bound) %>%
  { (.[-1] + .[-length(.)]) / 2 }

步骤2:使用自动截距绘图

将计算好的截距传入ggplot:

library(ggplot2)

ggplot(df, aes(x = price_index, y = quantity_index, label = products)) +
  coord_fixed() +
  geom_point(colour = 'blue', size = 3, alpha=0.9) +
  scale_x_continuous(expand = c(0, 0), limits = c(50, 100),
                     breaks = seq(50, 100, 5)) +
  scale_y_continuous(expand = c(0, 0), limits = c(50, 100),
                     breaks = seq(50, 100, 5)) +
  geom_vline(xintercept = price_cuts, color='black', linetype='dotted', size=1) +
  geom_hline(yintercept = quantity_cuts, color='black', linetype='dotted', size=1)

结果验证

计算出的截距与手动设定的目标值完全匹配:

  • price_cuts结果为 c(68.81026, 73.82950),对应需求的avg(70.50442,67.11609)和avg(74.41048,73.24853)
  • quantity_cuts结果为 c(77.57361, 90.38886),对应需求的avg(79.71225,75.43498)和avg(90.95081,89.82692)

内容的提问来源于stack exchange,提问作者ah bon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 12:07:04