如何自动设置散点图分割线截距,将散点划分为Top K层级
自动计算层级边界平均值作为ggplot截距
现有数据集df,其中price_index和quantity_index按层级划分,我用ggplot绘制散点图时手动设置了geom_vline和geom_hline的截距,现在需要自动计算相邻层级边界值的平均值作为截距——比如价格维度取avg(70.50442, 67.11609)和avg(74.41048, 73.24853)作为竖线截距,数量维度取avg(79.71225, 75.43498)和avg(90.95081, 89.82692)作为横线截距。
数据集预览
products price_index quantity_index price_rank quantity_rank 1 a 95.00000 95.00000 high_price high_volume 2 b 80.69012 94.53585 high_price high_volume 3 c 74.41048 90.95081 high_price high_volume 4 d 73.24853 89.82692 medium_price medium_volume 5 e 70.50442 79.71225 medium_price medium_volume 6 f 67.11609 75.43498 low_price low_volume 7 g 64.14685 58.26419 low_price low_volume 8 h 56.76375 56.16531 low_price low_volume 9 i 55.76472 56.02838 low_price low_volume 10 j 55.70475 50.24873 low_price low_volume
当前手动设置截距的绘图代码
ggplot(df, aes(x = price_index, y = quantity_index, label = products)) + coord_fixed() + geom_point(colour = 'blue', size = 3, alpha=0.9) + scale_x_continuous(expand = c(0, 0), limits = c(50, 100), breaks = seq(50, 100, 5) ) + scale_y_continuous(expand = c(0, 0), limits = c(50, 100), breaks = seq(50, 100, 5) ) + geom_vline(xintercept = c(68, 74), color='black', linetype='dotted', size=1) + geom_hline(yintercept = c(77, 90), color='black', linetype='dotted', size=1)
数据集结构
df <- structure(list(products = c("a", "b", "c", "d", "e", "f", "g", "h", "i", "j"), price_index = c(95, 80.69011538, 74.41047705, 73.24853055, 70.5044217, 67.1160916, 64.14685495, 56.76375355, 55.76472446, 55.70475052), quantity_index = c(95, 94.53585227, 90.95080999, 89.82692004, 79.71224701, 75.43498354, 58.26419203, 56.16530529, 56.02838119, 50.24873055), price_rank = c("high_price", "high_price", "high_price", "medium_price", "medium_price", "low_price", "low_price", "low_price", "low_price", "low_price"), quantity_rank = c("high_volume", "high_volume", "high_volume", "medium_volume", "medium_volume", "low_volume", "low_volume", "low_volume", "low_volume", "low_volume" )), class = "data.frame", row.names = c(NA, -10L))
解决方法
步骤1:自动计算截距
通过分组提取各层级的边界值,再计算相邻边界的平均值:
library(dplyr) # 计算价格维度截距:取各price_rank组的最小price_index,再算相邻组边界的平均值 price_cuts <- df %>% group_by(price_rank) %>% summarise(bound = min(price_index)) %>% arrange(desc(bound)) %>% # 按边界从高到低排序,确保层级对应正确 pull(bound) %>% { (.[-1] + .[-length(.)]) / 2 } # 相邻值相加取平均 # 计算数量维度截距:逻辑同上 quantity_cuts <- df %>% group_by(quantity_rank) %>% summarise(bound = min(quantity_index)) %>% arrange(desc(bound)) %>% pull(bound) %>% { (.[-1] + .[-length(.)]) / 2 }
步骤2:使用自动截距绘图
将计算好的截距传入ggplot:
library(ggplot2) ggplot(df, aes(x = price_index, y = quantity_index, label = products)) + coord_fixed() + geom_point(colour = 'blue', size = 3, alpha=0.9) + scale_x_continuous(expand = c(0, 0), limits = c(50, 100), breaks = seq(50, 100, 5)) + scale_y_continuous(expand = c(0, 0), limits = c(50, 100), breaks = seq(50, 100, 5)) + geom_vline(xintercept = price_cuts, color='black', linetype='dotted', size=1) + geom_hline(yintercept = quantity_cuts, color='black', linetype='dotted', size=1)
结果验证
计算出的截距与手动设定的目标值完全匹配:
price_cuts结果为c(68.81026, 73.82950),对应需求的avg(70.50442,67.11609)和avg(74.41048,73.24853)quantity_cuts结果为c(77.57361, 90.38886),对应需求的avg(79.71225,75.43498)和avg(90.95081,89.82692)
内容的提问来源于stack exchange,提问作者ah bon
相关产品推荐
相关产品推荐

