You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用更简洁的dplyr方法替换分组中中间位置的频率值

问题描述

现有按rowid分组的数据集,包含频率f、位置position和单词word字段,数据如下:

df
   rowid    word   f position
1      2       i 700        1
2      2      'm 600        2
3      2    fine   1        3
4      3     how 400        1
5      3      's 500        2
6      3     the 700        3
7      3 weather  20        4
8      4      it 390        1
9      4      's 500        2
10     4  really 177        3
11     4    very 200        4
12     4    cold  35        5
13     5       i 700        1
14     5    love 199        2
15     5     you 400        3

需求

对position数量超过3的rowid分组执行以下操作:

  • 将所有中间position的f值替换为这些中间值的平均值
  • 合并中间位置的word
  • 调整position标签(最终每组保留3个位置:1、中间合并后的2、末尾的3)

现有问题

已有基于dplyr的实现方案,但流程繁琐,且实际数据中使用left_join会出现问题,需要更简洁、无需left_join的dplyr实现方法。

数据集定义代码:

df <- data.frame(
  rowid = c(2,2,2,3,3,3,3,4,4,4,4,4,5,5,5),
  word = c("i","'m","fine",
           "how","'s","the","weather",
           "it","'s","really", "very","cold",
           "i","love","you"),
  f = c(700,600,1,
        400,500,700,20,
        390,500,177,200,35,
        700,199,400),
  position = c(1,2,3,
               1,2,3,4,
               1,2,3,4,5,
               1,2,3)
)

简洁dplyr实现方案

library(dplyr)

df_processed <- df %>%
  group_by(rowid) %>%
  mutate(
    # 标记首尾位置
    is_edge = position %in% c(1, max(position)),
    # 计算中间f值的平均值(仅针对position数>3的组)
    mid_f_mean = if(n() > 3) mean(f[!is_edge]) else NA,
    # 合并中间位置的word(仅针对position数>3的组)
    mid_word = if(n() > 3) paste(word[!is_edge], collapse = " ") else NA
  ) %>%
  # 保留每组的首尾行
  filter(is_edge) %>%
  # 为需要处理的组添加合并后的中间行
  bind_rows(
    df %>%
      group_by(rowid) %>%
      filter(n() > 3) %>%
      summarise(
        word = first(mid_word),
        f = first(mid_f_mean),
        position = 2,
        .groups = "drop"
      )
  ) %>%
  # 按rowid和position排序,保证结构正确
  arrange(rowid, position) %>%
  # 清理临时辅助列
  select(-is_edge, -mid_f_mean, -mid_word) %>%
  ungroup()

print(df_processed)

代码说明

  1. 分组预处理:按rowid分组后,标记首尾位置,同时计算中间f的平均值和合并后的word文本。
  2. 构建结果集:先保留每组的首尾行,再为position数量超过3的组单独生成合并后的中间行,通过bind_rows拼接在一起。
  3. 整理输出:按组和位置排序,移除临时辅助列,得到符合需求的最终数据。

输出结果

# A tibble: 12 × 4
   rowid word          f position
   <dbl> <chr>      <dbl>    <dbl>
 1     2 i            700        1
 2     2 'm           600        2
 3     2 fine           1        3
 4     3 how          400        1
 5     3 "'s the"     360        2
 6     3 weather       20        3
 7     4 it           390        1
 8     4 "'s really very"   225.        2
 9     4 cold          35        3
10     5 i            700        1
11     5 love         199        2
12     5 you          400        3

内容的提问来源于stack exchange,提问作者Chris Ruehlemann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 08:22:49