R语言:基于多匹配规则的data.table列名匹配代码优化问询
优化data.table按匹配规则查找列名的代码
问题背景
现有两个data.table对象:
DT_1存储待匹配字符串及匹配规则:
library(data.table) DT_1 <- data.table(Source_name = c("Apple","Banana","Orange","Pear","Random"), Match_type = c("Anywhere","Beginning","Anywhere","End","End"))
DT_2是目标数据集,需从它的列名中按规则匹配:
DT_2 <- data.table(Pear_1 = 1,Eat_apple = 1,Split_Banana = 1, Pear_2 = 1,Eat_pear = 1,Orange_peel = 1,Banana_chair = 1)
需求
根据DT_1中每行的Match_type(不区分大小写),找到DT_2列名中与Source_name首次匹配的列名:
Anywhere:字符串出现在列名任意位置Beginning:字符串匹配列名开头End:字符串匹配列名结尾
示例:
- "Apple"(Anywhere)的首个匹配列是
Eat_apple - "Banana"(Beginning)的首个匹配列是
Banana_chair
原始实现(繁琐版本)
已实现功能,但代码冗余:
library(purrr) DT_1[,Col_Name := names(DT_2)[unlist(pmap(.l = .SD[,.(x = Match_type,y = Source_name)], .f = function(x,y){ if(x == "Anywhere"){ grep(tolower(y),tolower(names(DT_2)))[1] # 返回任意位置的首个匹配项 }else if (x == "Beginning"){ grep(paste0("^",tolower(y),""),tolower(names(DT_2)))[1] # 返回开头匹配的首个项 }else if (x == "End"){ grep(paste0("",tolower(y),"$"),tolower(names(DT_2)))[1] # 返回结尾匹配的首个项 }}))]]
优化方案
方案1:简化正则生成+data.table原生操作
预生成匹配规则对应的正则表达式,用原生操作替代purrr,减少分支判断:
library(data.table) # 预存转小写的列名,避免重复转换 col_names_lower <- tolower(names(DT_2)) col_names <- names(DT_2) # 生成对应匹配规则的正则表达式 DT_1[, regex := fcase( Match_type == "Anywhere", tolower(Source_name), Match_type == "Beginning", paste0("^", tolower(Source_name)), Match_type == "End", paste0(tolower(Source_name), "$") )] # 逐行匹配并提取首个匹配列名 DT_1[, Col_Name := col_names[ sapply(regex, function(r) which(grepl(r, col_names_lower))[1]) ]] # 可选:删除临时生成的regex列 DT_1[, regex := NULL]
方案2:用stringr实现简洁匹配
针对长度不匹配问题,对每行单独处理,避免批量操作的冲突:
library(data.table) library(stringr) col_names_lower <- tolower(names(DT_2)) col_names <- names(DT_2) DT_1[, Col_Name := col_names[ sapply(1:.N, function(i) { y <- tolower(Source_name[i]) pattern <- switch(Match_type[i], Anywhere = y, Beginning = str_c("^", y), End = str_c(y, "$") ) str_which(col_names_lower, pattern)[1] }) ]]
方案3:向量化处理(高效版)
预先构建所有匹配逻辑的矩阵,通过索引直接提取结果,避免循环:
library(data.table) library(stringr) col_names_lower <- tolower(names(DT_2)) col_names <- names(DT_2) source_lower <- tolower(DT_1$Source_name) # 预计算所有匹配规则的结果位置 matches <- data.table( Anywhere = sapply(source_lower, function(s) str_which(col_names_lower, s)[1]), Beginning = sapply(source_lower, function(s) str_which(col_names_lower, str_c("^", s))[1]), End = sapply(source_lower, function(s) str_which(col_names_lower, str_c(s, "$"))[1]) ) # 根据Match_type提取对应结果 DT_1[, Col_Name := col_names[matches[cbind(1:.N, Match_type)]]]
验证结果
运行优化代码后,DT_1的Col_Name列结果如下:
| Source_name | Match_type | Col_Name |
|---|---|---|
| Apple | Anywhere | Eat_apple |
| Banana | Beginning | Banana_chair |
| Orange | Anywhere | Orange_peel |
| Pear | End | Pear_1 |
| Random | End | NA |
内容的提问来源于stack exchange,提问作者Phil
相关产品推荐
相关产品推荐

