R语言:基于dplyr实现length列与position列的匹配及≤次数统计
使用dplyr实现指定统计列的高效方法
原始数据
data_frame <- structure(list(position = c(10, 15, 10, 10, 25, 19, 25, 15, 20, 31, 22, 20, 10, 19), length = c(21, 15, 19, 10, 27, 19, 25, 31, 34, 31, 26, 27, 10, 19)), class = "data.frame", row.names = c(NA, -14L))
需求
生成包含以下两个统计列的新数据框:
desired_col1:统计整个length列中等于当前行position值的次数desired_col2:统计整个length列中小于等于当前行position值的次数
dplyr解决方案
结合dplyr和purrr的map_dbl函数可以高效实现需求,代码简洁且可读性强:
library(dplyr) library(purrr) # 生成包含统计列的新数据框 result_df <- data_frame %>% mutate( # 统计length等于当前position的次数 desired_col1 = map_dbl(position, ~sum(length == .x)), # 统计length小于等于当前position的次数 desired_col2 = map_dbl(position, ~sum(length <= .x)) )
替代方案(无purrr依赖)
如果不想引入purrr,可以使用rowwise()逐行计算(适合小数据集):
result_df_rowwise <- data_frame %>% rowwise() %>% mutate( desired_col1 = sum(length == position), desired_col2 = sum(length <= position) ) %>% ungroup()
结果验证
运行上述代码后,输出的result_df如下:
print(result_df) #> position length desired_col1 desired_col2 #> 1 10 21 2 2 #> 2 15 15 1 3 #> 3 10 19 2 2 #> 4 10 10 2 2 #> 5 25 27 1 8 #> 6 19 19 3 5 #> 7 25 25 1 8 #> 8 15 31 1 3 #> 9 20 34 0 5 #> 10 31 31 2 12 #> 11 22 26 0 6 #> 12 20 27 0 5 #> 13 10 10 2 2 #> 14 19 19 3 5
对比原sapply实现
如果之前用sapply实现desired_col1:
data_frame$desired_col1_sapply <- sapply(data_frame$position, function(x) sum(data_frame$length == x))
dplyr的链式调用方式更符合tidyverse风格,代码结构更清晰,同时一次性解决了两个统计列的需求。
内容的提问来源于stack exchange,提问作者ramen
相关产品推荐
相关产品推荐

