You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何替换cumsum实现R数据框分组ID的指定后缀赋值?

R数据框分组添加指定后缀的解决方案

问题背景

现有如下R数据框:

df1 <- structure(list(id = c("2023021112", "2023021112", "2023021112", 
"2023021112", "2023021112", "2023021112", "2023021112", "2023021112", 
"2023021112", "2023021112", "2023021113", "2023021113", "2023021113", 
"2023021113", "2023021113", "2023021113", "2023021113", "2023021113", 
"2023021113", "2023021113"), response = c("1", "Happy", "Sad", 
"Neutral", "Fearful", "2", "Disgusted", "Happy", "Sad", "Surprised", "2", "Sad", "Sad", 
"Neutral", "Fearful", "1", "Disgusted", "Happy", "Sad", "Fearful"
)), row.names = c(1L, 2L, 3L, 4L, 5L, 72L, 73L, 74L, 75L, 76L, 6L, 7L, 8L, 9L, 10L, 77L, 78L, 79L, 80L, 91L), class = "data.frame")

需求如下:

  • 将response列中数字行到下一个数字行前的所有行,为其id添加-0n后缀(n为对应数字行的数字,例如response为"1"的行后续行对应2023021112-01)
  • 删除response列包含数字的行,最终得到目标数据框:
structure(list(id = c("2023021112-01", "2023021112-01", 
"2023021112-01", "2023021112-01", "2023021112-02", "2023021112-02", "2023021112-02", 
"2023021112-02", "2023021112-02", "2023021113-02", "2023021113-02", "2023021113-02", 
"2023021113-01", "2023021113-01", "2023021113-01", 
"2023021113-01"), response = c("Happy", "Sad", 
"Neutral", "Fearful", "Disgusted", "Happy", "Sad", "Surprised", "Sad", "Sad", 
"Neutral", "Fearful", "Disgusted", "Happy", "Sad", "Fearful"
)), row.names = c(2L, 3L, 4L, 5L, 73L, 74L, 75L, 76L, 7L, 8L, 9L, 10L, 78L, 79L, 80L, 91L), class = "data.frame")

此前使用cumsum的方案会生成全局连续的后缀(-01、-02、-03、-04),无法满足按每个id内数字分组的需求。

解决方案

核心思路是按id分组后,提取response中的数字作为分组标识,向下填充给后续非数字行,再用该数字生成后缀。具体代码如下:

library(dplyr)
library(tidyr)

result <- df1 %>%
  # 标记数字行的分组编号,非数字行设为NA
  mutate(group_num = ifelse(grepl("^\\d+$", response), as.numeric(response), NA)) %>%
  # 按id分组,向下填充分组编号,让非数字行继承最近的数字
  group_by(id) %>%
  fill(group_num, .direction = "down") %>%
  # 拼接id和格式化后的后缀(两位数字)
  mutate(id = paste(id, sprintf("%02d", group_num), sep = "-")) %>%
  ungroup() %>%
  # 过滤掉response为数字的行
  filter(!grepl("^\\d+$", response)) %>%
  # 移除临时列
  select(-group_num)

# 输出结果
print(result)

代码关键点说明

  • grepl("^\\d+$", response):精准匹配纯数字的response行
  • fill(group_num, .direction = "down"):在每个id分组内,将数字行的编号向下填充,确保后续非数字行对应正确的后缀
  • sprintf("%02d", group_num):将数字格式化为两位带前导零的字符串,满足-0n的后缀格式

内容的提问来源于stack exchange,提问作者grace.cutler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 01:54:52