如何用更简洁的dplyr方法替换分组中中间位置的频率值
问题描述
现有按rowid分组的数据集,包含频率f、位置position和单词word字段,数据如下:
df rowid word f position 1 2 i 700 1 2 2 'm 600 2 3 2 fine 1 3 4 3 how 400 1 5 3 's 500 2 6 3 the 700 3 7 3 weather 20 4 8 4 it 390 1 9 4 's 500 2 10 4 really 177 3 11 4 very 200 4 12 4 cold 35 5 13 5 i 700 1 14 5 love 199 2 15 5 you 400 3
需求
对position数量超过3的rowid分组执行以下操作:
- 将所有中间
position的f值替换为这些中间值的平均值 - 合并中间位置的
word - 调整
position标签(最终每组保留3个位置:1、中间合并后的2、末尾的3)
现有问题
已有基于dplyr的实现方案,但流程繁琐,且实际数据中使用left_join会出现问题,需要更简洁、无需left_join的dplyr实现方法。
数据集定义代码:
df <- data.frame( rowid = c(2,2,2,3,3,3,3,4,4,4,4,4,5,5,5), word = c("i","'m","fine", "how","'s","the","weather", "it","'s","really", "very","cold", "i","love","you"), f = c(700,600,1, 400,500,700,20, 390,500,177,200,35, 700,199,400), position = c(1,2,3, 1,2,3,4, 1,2,3,4,5, 1,2,3) )
简洁dplyr实现方案
library(dplyr) df_processed <- df %>% group_by(rowid) %>% mutate( # 标记首尾位置 is_edge = position %in% c(1, max(position)), # 计算中间f值的平均值(仅针对position数>3的组) mid_f_mean = if(n() > 3) mean(f[!is_edge]) else NA, # 合并中间位置的word(仅针对position数>3的组) mid_word = if(n() > 3) paste(word[!is_edge], collapse = " ") else NA ) %>% # 保留每组的首尾行 filter(is_edge) %>% # 为需要处理的组添加合并后的中间行 bind_rows( df %>% group_by(rowid) %>% filter(n() > 3) %>% summarise( word = first(mid_word), f = first(mid_f_mean), position = 2, .groups = "drop" ) ) %>% # 按rowid和position排序,保证结构正确 arrange(rowid, position) %>% # 清理临时辅助列 select(-is_edge, -mid_f_mean, -mid_word) %>% ungroup() print(df_processed)
代码说明
- 分组预处理:按
rowid分组后,标记首尾位置,同时计算中间f的平均值和合并后的word文本。 - 构建结果集:先保留每组的首尾行,再为
position数量超过3的组单独生成合并后的中间行,通过bind_rows拼接在一起。 - 整理输出:按组和位置排序,移除临时辅助列,得到符合需求的最终数据。
输出结果
# A tibble: 12 × 4 rowid word f position <dbl> <chr> <dbl> <dbl> 1 2 i 700 1 2 2 'm 600 2 3 2 fine 1 3 4 3 how 400 1 5 3 "'s the" 360 2 6 3 weather 20 3 7 4 it 390 1 8 4 "'s really very" 225. 2 9 4 cold 35 3 10 5 i 700 1 11 5 love 199 2 12 5 you 400 3
内容的提问来源于stack exchange,提问作者Chris Ruehlemann
相关产品推荐
相关产品推荐

