在R中基于向量匹配筛选文本:提取B中未出现于A的元素
R实现术语匹配与筛选需求
原始数据
A <- data.frame(c("Absolute Value", "absolute deviation", "acceptance line ; acceptance boundary", "age-adjusted rate", "variance", "modified mean ; modified arithmetic mean ; trimmed mean ")) B <- data.frame(c("descriptive", "Acceptance Boundary", "deviation", "modified arithmetic mean", "mutability ; variability"))
需求说明
生成数据框C,包含B中未出现在A中的术语,需满足:
- 忽略大小写差异(如
Acceptance Boundary与acceptance boundary视为同一术语) - 识别分号分隔的多术语(如A中的
acceptance line ; acceptance boundary包含acceptance boundary术语)
预期结果:
C <- data.frame(c("descriptive", "deviation", "mutability ; variability"))
解决方案代码
# 安装并加载stringr包(首次使用需执行安装命令) # install.packages("stringr") library(stringr) # 提取A中所有独立术语(统一小写、去除前后空格) A_terms <- str_split(A[, 1], ";") %>% unlist() %>% str_trim() %>% tolower() # 筛选B中完全未出现在A里的条目 C <- B[sapply(B[, 1], function(x) { # 标准化处理当前B条目内的术语 b_split <- str_split(x, ";") %>% unlist() %>% str_trim() %>% tolower() # 检查该条目下所有术语是否都不在A的术语集合中 all(!b_split %in% A_terms) }), , drop = FALSE] # 输出结果 print(C)
代码解释
- 处理A数据:拆分所有分号分隔的内容,去除术语前后空格并统一转为小写,得到A包含的所有独立术语集合。
- 筛选B数据:遍历B的每个条目,对条目内的术语做同样的标准化处理,检查该条目下所有术语是否都未出现在A的集合中,符合条件的条目保留到C。
- 最终输出的C即为符合需求的结果。
内容的提问来源于stack exchange,提问作者nickolakis
相关产品推荐
相关产品推荐

