如何将data.table中指定序列的多行替换为单行?
没问题,我来帮你搞定这个序列匹配和替换的需求!下面我会分别给出data.table和data.frame两种实现方案,完全贴合你的预期结果。
基于data.table的解决方案
首先先确认原始数据:
library(data.table) DT <- data.table(id=c(1,1,1,1,2,2,2,2), place = c("a","b","c","d","a","b","d","e"), seq = c(1,2,3,4,1,2,3,4)) setorder(DT, id, seq)
核心思路
针对每个id分组,先识别出连续的"b","c","d"序列,标记出需要替换的"b"和"c"行,然后移除这两行,添加一行place="z"的记录(沿用"b"行的seq值),最后重新排序。
具体代码
# 按id分组,获取下一个、下下个位置的place值,用于匹配序列 DT[, `:=`( next_place = shift(place, type = "lead"), next_next_place = shift(place, n = 2, type = "lead") ), by = id] # 标记出属于目标序列的b行(后面跟着c和d) DT[, is_target_b := place == "b" & next_place == "c" & next_next_place == "d"] # 标记出属于目标序列的c行(前面是b,后面是d) DT[, is_target_c := place == "c" & shift(place) == "b" & next_place == "d"] # 生成最终结果:保留非目标行 + 添加替换后的z行 result <- rbind( # 保留不需要替换的行(排除目标b和c) DT[!is_target_b & !is_target_c, .(id, place, seq)], # 添加替换后的z行,取目标b行的信息 DT[is_target_b, .(id, place = "z", seq)] ) # 重新按id和seq排序,得到预期格式 setorder(result, id, seq) print(result)
运行后输出和你给出的DT.tobe完全一致:
id place seq 1: 1 a 1 2: 1 z 2 3: 1 d 4 4: 2 a 1 5: 2 b 2 6: 2 d 3 7: 2 e 4
基于data.frame(dplyr)的解决方案
如果你习惯用data.frame和dplyr语法,也可以用下面的实现:
library(dplyr) DT_df <- as.data.frame(DT) result_df <- DT_df %>% group_by(id) %>% # 标记目标序列的行 mutate( next_place = lead(place), next_next_place = lead(place, n = 2), is_target_b = place == "b" & next_place == "c" & next_next_place == "d", is_target_c = place == "c" & lag(place) == "b" & next_place == "d" ) %>% # 保留非目标行 filter(!is_target_b & !is_target_c) %>% # 添加替换后的z行 bind_rows( DT_df %>% group_by(id) %>% filter(place == "b" & lead(place) == "c" & lead(place, n=2) == "d") %>% mutate(place = "z") %>% select(id, place, seq) ) %>% # 重新排序 arrange(id, seq) %>% ungroup() print(result_df)
这个方案的逻辑和data.table版本完全一致,输出结果也相同。
内容的提问来源于stack exchange,提问作者User800701
相关产品推荐
相关产品推荐

