如何基于向量中相似(非完全匹配)元素筛选DataFrame子集?
问题描述
我有一个2914×6的DataFrame,其中一列是类似bird_F.pw的字符串(前缀为动物类群,后缀为物种缩写)。需要根据单独的物种缩写向量(如F.pw),提取该列中后缀匹配这些缩写的所有行。尝试过%in%、%like%运算符,但无法实现非完全匹配的筛选。
示例数据
构建示例DataFrame
df <- cbind( c("A","B","C","D","E"), c(1:5), c("insect_F.vp","bird_L.ts","insect_P.qr","insect_V.cl","bird_H.dw") ) colnames(df) <- c("season","survey_id","pollinator") # 转换为DataFrame(cbind默认生成矩阵) df <- as.data.frame(df)
待搜索的物种缩写向量
abbrevs <- c("L.ts","P.qr","H.dw")
预期输出
output <- cbind(c("B","C","E"),c(2:3,5),c("bird_L.ts","insect_P.qr","bird_H.dw")) colnames(output) <- colnames(df) output <- as.data.frame(output)
解决方案
方法1:基础R实现(正则匹配)
将缩写向量拼接成正则表达式,匹配字符串末尾的目标缩写:
# 构建正则模式:匹配任意前缀后接目标缩写,且缩写位于字符串结尾 pattern <- paste0("(", paste(abbrevs, collapse = "|"), ")$") # 筛选符合条件的行 result <- df[grepl(pattern, df$pollinator), ]
方法2:tidyverse工具链实现
用dplyr的筛选函数结合stringr的字符串检测功能:
library(dplyr) library(stringr) result <- df %>% filter(str_detect(pollinator, paste0("(", paste(abbrevs, collapse = "|"), ")$")))
方法3:拆分字符串后匹配
如果前缀和缩写始终以下划线分隔,可以拆分字符串提取后缀后再匹配:
# 拆分pollinator列,提取物种缩写后缀 df$species_abbrev <- sapply(strsplit(df$pollinator, "_"), "[", 2) # 筛选后缀在目标向量中的行 result <- df[df$species_abbrev %in% abbrevs, ] # 可选:删除临时生成的后缀列 result <- result[, !colnames(result) %in% "species_abbrev"]
内容的提问来源于stack exchange,提问作者ElizaBeso000
相关产品推荐
相关产品推荐

