You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:基于多匹配规则的data.table列名匹配代码优化问询

优化data.table按匹配规则查找列名的代码

问题背景

现有两个data.table对象:

  • DT_1存储待匹配字符串及匹配规则:
library(data.table)  
DT_1 <- data.table(Source_name = c("Apple","Banana","Orange","Pear","Random"),
                   Match_type = c("Anywhere","Beginning","Anywhere","End","End"))
  • DT_2是目标数据集,需从它的列名中按规则匹配:
DT_2 <- data.table(Pear_1 = 1,Eat_apple = 1,Split_Banana = 1,
                   Pear_2 = 1,Eat_pear = 1,Orange_peel = 1,Banana_chair = 1)

需求

根据DT_1中每行的Match_type(不区分大小写),找到DT_2列名中与Source_name首次匹配的列名:

  • Anywhere:字符串出现在列名任意位置
  • Beginning:字符串匹配列名开头
  • End:字符串匹配列名结尾

示例:

  • "Apple"(Anywhere)的首个匹配列是Eat_apple
  • "Banana"(Beginning)的首个匹配列是Banana_chair

原始实现(繁琐版本)

已实现功能,但代码冗余:

library(purrr)      
DT_1[,Col_Name := names(DT_2)[unlist(pmap(.l = .SD[,.(x = Match_type,y = Source_name)],
              .f = function(x,y){
                  if(x == "Anywhere"){
                       grep(tolower(y),tolower(names(DT_2)))[1] # 返回任意位置的首个匹配项
                  }else if (x == "Beginning"){
                       grep(paste0("^",tolower(y),""),tolower(names(DT_2)))[1] # 返回开头匹配的首个项
                  }else if (x == "End"){
                        grep(paste0("",tolower(y),"$"),tolower(names(DT_2)))[1] # 返回结尾匹配的首个项
                  }}))]]

优化方案

方案1:简化正则生成+data.table原生操作

预生成匹配规则对应的正则表达式,用原生操作替代purrr,减少分支判断:

library(data.table)

# 预存转小写的列名,避免重复转换
col_names_lower <- tolower(names(DT_2))
col_names <- names(DT_2)

# 生成对应匹配规则的正则表达式
DT_1[, regex := fcase(
  Match_type == "Anywhere", tolower(Source_name),
  Match_type == "Beginning", paste0("^", tolower(Source_name)),
  Match_type == "End", paste0(tolower(Source_name), "$")
)]

# 逐行匹配并提取首个匹配列名
DT_1[, Col_Name := col_names[
  sapply(regex, function(r) which(grepl(r, col_names_lower))[1])
]]

# 可选:删除临时生成的regex列
DT_1[, regex := NULL]

方案2:用stringr实现简洁匹配

针对长度不匹配问题,对每行单独处理,避免批量操作的冲突:

library(data.table)
library(stringr)

col_names_lower <- tolower(names(DT_2))
col_names <- names(DT_2)

DT_1[, Col_Name := col_names[
  sapply(1:.N, function(i) {
    y <- tolower(Source_name[i])
    pattern <- switch(Match_type[i],
      Anywhere = y,
      Beginning = str_c("^", y),
      End = str_c(y, "$")
    )
    str_which(col_names_lower, pattern)[1]
  })
]]

方案3:向量化处理(高效版)

预先构建所有匹配逻辑的矩阵,通过索引直接提取结果,避免循环:

library(data.table)
library(stringr)

col_names_lower <- tolower(names(DT_2))
col_names <- names(DT_2)
source_lower <- tolower(DT_1$Source_name)

# 预计算所有匹配规则的结果位置
matches <- data.table(
  Anywhere = sapply(source_lower, function(s) str_which(col_names_lower, s)[1]),
  Beginning = sapply(source_lower, function(s) str_which(col_names_lower, str_c("^", s))[1]),
  End = sapply(source_lower, function(s) str_which(col_names_lower, str_c(s, "$"))[1])
)

# 根据Match_type提取对应结果
DT_1[, Col_Name := col_names[matches[cbind(1:.N, Match_type)]]]

验证结果

运行优化代码后,DT_1的Col_Name列结果如下:

Source_nameMatch_typeCol_Name
AppleAnywhereEat_apple
BananaBeginningBanana_chair
OrangeAnywhereOrange_peel
PearEndPear_1
RandomEndNA

内容的提问来源于stack exchange,提问作者Phil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 05:15:22