R语言:分组处理字符串,移除长串中非目标匹配子串
R语言实现:按组移除长字符串中除目标子串外的其他同组子串
需求说明
现有一个R语言数据框,其中V1列为长字符串,每行包含至少两个可匹配的子串;V2列为该行需要保留的目标子串。要求:对每组相同V1值的行,将该行V1中除对应V2之外的同组其他V2子串移除,仅保留包含目标V2的长串部分。
示例输入
library(qdap) library(magrittr) t <- read.table(text=" V1,V2 Video of all presentations and discussionPart 1Part 2Video,Part 1 Video of all presentations and discussionPart 1Part 2Video,Part 2 Video of all presentations and discussionPart 1Part 2Video,Video Background - PDFVideo Soil management and update (pdf)Video,PDF Background - PDFVideo Soil management and update (pdf)Video,Video Background - PDFVideo Soil management and update (pdf)Video,Soil management and update (pdf) Background - PDFVideo Soil management and update (pdf)Video,Video ", header=T, sep = ",")
尝试的无效代码
t%>% split(.,.$V1)%>% lapply(.,function(x){(unique(x$V2))})%>% lapply(.,function(y){mgsub(pattern=y[[1]],replacement="",names(y))})
期望输出
t <- read.table(text=" V1,V2 Video of all presentations and discussionPart 1,Part 1 Video of all presentations and discussionPart 2,Part 2 Video of all presentations and discussionVideo,Video Background - PDF,PDF Background - Video,Video Background - Soil management and update (pdf),Soil management and update (pdf) Background - Video,Video ", header=T, sep = ",")
解决方案
使用dplyr分组处理,结合stringr的字符串操作实现需求:
library(dplyr) library(stringr) # 读取示例数据(若已加载可跳过) t <- read.table(text=" V1,V2 Video of all presentations and discussionPart 1Part 2Video,Part 1 Video of all presentations and discussionPart 1Part 2Video,Part 2 Video of all presentations and discussionPart 1Part 2Video,Video Background - PDFVideo Soil management and update (pdf)Video,PDF Background - PDFVideo Soil management and update (pdf)Video,Video Background - PDFVideo Soil management and update (pdf)Video,Soil management and update (pdf) Background - PDFVideo Soil management and update (pdf)Video,Video ", header=T, sep=",", stringsAsFactors=F) # 核心处理逻辑 result <- t %>% group_by(V1) %>% mutate( # 生成当前行需要移除的子串正则:组内唯一V2排除当前行V2 remove_pattern = str_c(setdiff(unique(V2), V2), collapse = "|"), # 移除匹配子串并清理多余空格 new_V1 = str_remove_all(V1, remove_pattern) %>% str_squish() ) %>% ungroup() %>% # 替换原V1列并整理输出 select(new_V1, V2) %>% rename(V1 = new_V1) # 输出结果 print(result)
代码说明
- 分组:按
V1分组,确保仅处理同组内的子串; - 生成移除规则:用
setdiff筛选出当前组内除当前行V2外的所有唯一子串,拼接成正则表达式; - 移除子串:用
str_remove_all批量移除匹配的子串,str_squish清理可能出现的多余空格; - 整理输出:替换原
V1列并取消分组,得到符合要求的结果。
内容的提问来源于stack exchange,提问作者CrunchyTopping
相关产品推荐
相关产品推荐

