You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:分组处理字符串,移除长串中非目标匹配子串

R语言实现:按组移除长字符串中除目标子串外的其他同组子串

需求说明

现有一个R语言数据框,其中V1列为长字符串,每行包含至少两个可匹配的子串;V2列为该行需要保留的目标子串。要求:对每组相同V1值的行,将该行V1中除对应V2之外的同组其他V2子串移除,仅保留包含目标V2的长串部分。

示例输入

library(qdap)
library(magrittr)
t <- read.table(text="
V1,V2
  Video of all presentations and discussionPart 1Part 2Video,Part 1
  Video of all presentations and discussionPart 1Part 2Video,Part 2
  Video of all presentations and discussionPart 1Part 2Video,Video
  Background - PDFVideo Soil management and update (pdf)Video,PDF
  Background - PDFVideo Soil management and update (pdf)Video,Video
  Background - PDFVideo Soil management and update (pdf)Video,Soil management and update (pdf)
  Background - PDFVideo Soil management and update (pdf)Video,Video
", header=T, sep = ",")

尝试的无效代码

t%>%  
  split(.,.$V1)%>% 
    lapply(.,function(x){(unique(x$V2))})%>%
      lapply(.,function(y){mgsub(pattern=y[[1]],replacement="",names(y))})

期望输出

t <- read.table(text="
V1,V2
Video of all presentations and discussionPart 1,Part 1
Video of all presentations and discussionPart 2,Part 2
Video of all presentations and discussionVideo,Video
Background - PDF,PDF
Background - Video,Video
Background - Soil management and update (pdf),Soil management and update (pdf)
Background - Video,Video
", header=T, sep = ",")

解决方案

使用dplyr分组处理,结合stringr的字符串操作实现需求:

library(dplyr)
library(stringr)

# 读取示例数据(若已加载可跳过)
t <- read.table(text="
V1,V2
  Video of all presentations and discussionPart 1Part 2Video,Part 1
  Video of all presentations and discussionPart 1Part 2Video,Part 2
  Video of all presentations and discussionPart 1Part 2Video,Video
  Background - PDFVideo Soil management and update (pdf)Video,PDF
  Background - PDFVideo Soil management and update (pdf)Video,Video
  Background - PDFVideo Soil management and update (pdf)Video,Soil management and update (pdf)
  Background - PDFVideo Soil management and update (pdf)Video,Video
", header=T, sep=",", stringsAsFactors=F)

# 核心处理逻辑
result <- t %>%
  group_by(V1) %>%
  mutate(
    # 生成当前行需要移除的子串正则:组内唯一V2排除当前行V2
    remove_pattern = str_c(setdiff(unique(V2), V2), collapse = "|"),
    # 移除匹配子串并清理多余空格
    new_V1 = str_remove_all(V1, remove_pattern) %>% str_squish()
  ) %>%
  ungroup() %>%
  # 替换原V1列并整理输出
  select(new_V1, V2) %>%
  rename(V1 = new_V1)

# 输出结果
print(result)

代码说明

  1. 分组:按V1分组,确保仅处理同组内的子串;
  2. 生成移除规则:用setdiff筛选出当前组内除当前行V2外的所有唯一子串,拼接成正则表达式;
  3. 移除子串:用str_remove_all批量移除匹配的子串,str_squish清理可能出现的多余空格;
  4. 整理输出:替换原V1列并取消分组,得到符合要求的结果。

内容的提问来源于stack exchange,提问作者CrunchyTopping

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 19:20:33