dplyr:基于多个最近值的分组内筛选问题
按分组匹配向量中每个元素的最近值问题
我有一个向量用于筛选大型数据集:
sepals=c(4:7)
精确匹配的实现
精确匹配筛选很简单,直接用%in%就能完成:
iris %>% filter(Sepal.Length %in% sepals)
运行结果:
Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <fct> 1 5 3.6 1.4 0.2 setosa 2 5 3.4 1.5 0.2 setosa 3 5 3 1.6 0.2 setosa 4 5 3.4 1.6 0.4 setosa 5 5 3.2 1.2 0.2 setosa 6 5 3.5 1.3 0.3 setosa 7 5 3.5 1.6 0.6 setosa 8 5 3.3 1.4 0.2 setosa 9 7 3.2 4.7 1.4 versicolor 10 5 2 3.5 1 versicolor 11 6 2.2 4 1 versicolor 12 6 2.9 4.5 1.5 versicolor
单个元素匹配最近值的实现
但我实际需要按Species分组,匹配sepals向量中每个元素的最近值。单个元素处理时能正常实现,比如匹配4:
iris %>% group_by(Species) %>% filter(abs(Sepal.Length-4)==min(abs(Sepal.Length-4)))
运行结果:
Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <fct> 1 4.3 3 1.1 0.1 setosa 2 4.9 2.4 3.3 1 versicolor 3 4.9 2.5 4.5 1.7 virginica
也可以用更简洁的写法:
iris %>% group_by(Species) %>% slice(which.min(abs(Sepal.Length-4)))
或者:
iris %>% group_by(Species) %>% slice_min(abs(Sepal.Length-4))
批量处理整个向量遇到的问题
但当我尝试对整个sepals向量执行操作时,会得到错误结果和警告:
iris %>% group_by(Species) %>% slice(which.min(abs(Sepal.Length-sepals)))
运行结果:
Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <fct> 1 5 3 1.6 0.2 setosa 2 5.2 2.7 3.9 1.4 versicolor 3 6 3 4.8 1.8 virginica
警告信息:
1: In Sepal.Length - sepals : longer object length is not a multiple of shorter object length
类似的问题也出现在以下代码中:
iris %>% group_by(Species) %>% slice_min(abs(Sepal.Length-sepals))
iris %>% group_by(Species) %>% filter(abs(Sepal.Length-sepals)==min(abs(Sepal.Length-sepals)))
解决方案
问题出在直接用向量做减法时,R会按循环规则处理长度不匹配的向量,导致计算逻辑错误。要实现对sepals中每个元素,按Species分组找最近值,需要先把sepals转换为数据框,和iris做交叉连接,再计算每个分组(Species+目标值)的最小差值:
library(dplyr) library(tidyr) # 将sepals转为数据框 sepals_df <- tibble(target_sepal = 4:7) # 交叉连接+分组计算 iris %>% cross_join(sepals_df) %>% group_by(Species, target_sepal) %>% slice_min(abs(Sepal.Length - target_sepal)) %>% ungroup()
运行结果会包含每个物种对应sepals中4、5、6、7的最近匹配行:
Sepal.Length Sepal.Width Petal.Length Petal.Width Species target_sepal <dbl> <dbl> <dbl> <dbl> <fct> <int> 1 4.3 3 1.1 0.1 setosa 4 2 5 3.6 1.4 0.2 setosa 5 3 5.8 4 1.2 0.2 setosa 6 4 5.8 4 1.2 0.2 setosa 7 5 4.9 2.4 3.3 1 versicolor 4 6 5 2 3.5 1 versicolor 5 7 6 2.2 4 1 versicolor 6 8 7 3.2 4.7 1.4 versicolor 7 9 4.9 2.5 4.5 1.7 virginica 4 10 5.6 2.8 4.9 2 virginica 5 11 6 3 4.8 1.8 virginica 6 12 7.2 3.6 6.1 2.5 virginica 7
如果某个组有多个行和目标值的差值相同,slice_min会返回所有匹配行,若只想返回一行,可以加上参数n=1:
iris %>% cross_join(sepals_df) %>% group_by(Species, target_sepal) %>% slice_min(abs(Sepal.Length - target_sepal), n=1) %>% ungroup()
内容的提问来源于stack exchange,提问作者frostygoat
相关产品推荐
相关产品推荐

