You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按行匹配其他列值实现字符串列部分内容替换删除

解决方案

纯data.table原生实现(推荐)

直接通过by = 1:nrow(dt)实现逐行处理,每次调用当前行的state、city值构造匹配规则,不会出现跨行误删的问题:

library(data.table)
# 测试数据
dt<- data.table(city = c("arecibo","arecibo","cabo rojo", "new york", "dallas"), 
state=c("pr", "pr", "pr", "ny", "tx"), 
name=c("frutas of pr arecibo", "dallas frutas of pr", "cabo rojo metal plant", "greens new york", "cowboy shoes dallas tx"), 
desired=c("frutas", "dallas frutas", "metal plant", "greens", "cowboy shoes"))

# 逐行处理替换
dt[, processed := {
  # 构造正则:匹配单词边界的of、当前行state、当前行city
  pattern <- paste0("\\b(", paste(c("of", state, city), collapse = "|"), ")\\b")
  # 替换后清除多余空格
  trimws(gsub(pattern, "", name))
}, by = 1:nrow(dt)]

运行后dt$processed和你给出的期望desired列完全一致。

tidyverse风格实现(习惯用dplyr可参考)

通过rowwise()实现逐行处理,逻辑和上面一致:

library(dplyr)
library(stringr)

dt <- dt %>% 
  rowwise() %>% 
  mutate(processed = str_remove_all(name, paste0("\\b(", paste(c("of", state, city), collapse = "|"), ")\\b")) %>% str_squish()) %>% 
  ungroup()
核心逻辑说明
  • 逐行处理保证每次构造匹配规则时,只用当前行的state、city值,不会把其他行的城市/州名加入匹配,彻底解决了全量城市向量导致的第二行dallas被误删的问题
  • 正则里的\\b是单词边界标识,避免短的州名/城市名误删其他单词的子串(比如州名pr不会误删price里的pr)
  • 最后的trimws/str_squish用于清除替换后残留的首尾空格、中间连续空格,保证输出结果格式干净

内容的提问来源于stack exchange,提问作者Magasinus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 00:45:02