R语言:如何从数据框URL列中仅保留指定单词?
解决方案
可以用两种方法实现需求,分别基于tidyverse工具包和基础R:
方法一:使用tidyverse(推荐)
借助stringr包的字符串提取函数,结合coalesce处理未匹配的情况:
library(tidyverse) # 原始数据框 df <- data.frame (URL = c("export-180100-from-ec-to-us", "import-420340-to-ir-from-es","export","Product"), X = c(100,200,50,600)) # 处理URL列 df <- df %>% mutate(URL = coalesce(str_extract(URL, "export|import"), URL)) # 查看结果 print(df)
输出结果:
URL X 1 export 100 2 import 200 3 export 50 4 Product 600
说明:
str_extract(URL, "export|import"):从URL字符串中提取"export"或"import",无匹配时返回NAcoalesce(...):优先取第一个非NA值,即匹配到目标词时用提取结果,否则保留原字符串
方法二:基础R实现
无需额外安装包,用grepl和sub组合完成:
# 原始数据框 df <- data.frame (URL = c("export-180100-from-ec-to-us", "import-420340-to-ir-from-es","export","Product"), X = c(100,200,50,600)) # 处理URL列 df$URL <- ifelse( grepl("export|import", df$URL), sub(".*(export|import).*", "\\1", df$URL), df$URL ) # 查看结果 print(df)
输出结果与方法一完全一致。
说明:
grepl("export|import", df$URL):判断每个URL是否包含目标词sub(".*(export|import).*", "\\1", df$URL):将包含目标词的字符串替换为匹配到的"export"或"import"(通过分组捕获实现)- 不满足条件时直接保留原URL
内容的提问来源于stack exchange,提问作者John Owen
相关产品推荐
相关产品推荐

