如何基于多列正则匹配生成新二进制列?含报错与优化需求
问题与解决方案
第一部分:原代码报错修复
原问题代码与错误
用户尝试生成匹配正则模式的二进制列时,使用rowwise()+mutate()报错,原代码及错误如下:
df <- data.frame( idx = 1:5, column_b = letters[1:5], column_c = c('abc', 'abc', 'def', 'def', 'ghi'), column_d = c('def', 'def', 'def', 'def', 'def'), column_e = c('ghi', 'ghi', 'ghi', 'abc', 'ghi') ) apply_factor <- function(df, factor, col_low, col_high, pattern) { df %>% rowwise() %>% mutate(factor = sum(c_across(as.data.frame(sapply(select(df, {{col_low}}:{{col_high}}), grepl, pattern={{pattern}})))), na.rm = TRUE) } apply_factor(df, factor = 'abc', 'column_c', 'column_e', pattern = "^abc")
报错信息:
Error in `mutate()`: ! Problem while computing `factor = sum(...)`. i The error occurred in row 1. Caused by error in `as_indices_impl()`: ! Must subset columns with a valid subscript vector. x Subscript has the wrong type `data.frame< column_c: logical column_d: logical column_e: logical >`. i It must be numeric or character.
报错原因
c_across()的作用是在rowwise()上下文里选取当前行的指定列,它需要的是列选择器(如列名范围、tidyselect语法),而非提前生成的完整逻辑值dataframe。原代码中把sapply()生成的逻辑表传给c_across(),违背了它的使用逻辑,导致报错。
修复后的代码
方案1:dplyr向量化实现(推荐,避免rowwise)
直接对指定列批量应用grepl,返回逻辑矩阵后用rowSums判断每行是否有匹配:
library(dplyr) apply_factor <- function(df, factor, col_low, col_high, pattern) { df %>% mutate({{factor}} := as.integer( rowSums(grepl(pattern, select(., {{col_low}}:{{col_high}})), na.rm = TRUE) > 0 )) } # 测试调用 result <- apply_factor(df, factor = 'abc', 'column_c', 'column_e', pattern = "^abc") print(result)
方案2:修正rowwise写法
如果一定要用rowwise(),需直接在c_across()中指定列范围,再对当前行的列值应用grepl:
apply_factor_rowwise <- function(df, factor, col_low, col_high, pattern) { df %>% rowwise() %>% mutate({{factor}} := as.integer( any(grepl(pattern, c_across({{col_low}}:{{col_high}})), na.rm = TRUE) )) %>% ungroup() # 必须取消rowwise,避免后续操作效率低下 } # 测试调用 result_rowwise <- apply_factor_rowwise(df, factor = 'abc', 'column_c', 'column_e', pattern = "^abc") print(result_rowwise)
第二部分:大数据集高效实现
当前方法的效率问题
原代码使用rowwise()+sapply()的组合,本质是逐行处理,在百万行、几十列的数据集上会非常缓慢——rowwise()会破坏向量化优势,导致性能骤降。
高效实现方案
方案1:dplyr向量化(适合中等规模大数据)
前面的rowSums方案已经是向量化操作,比rowwise()快几个数量级,因为它直接对整个矩阵进行运算,无需逐行循环。
方案2:data.table(超大规模数据首选)
data.table的向量化运算性能远超基础dplyr,适合百万行以上的数据集:
library(data.table) # 转换为data.table格式 setDT(df) apply_factor_dt <- function(dt, factor, col_low, col_high, pattern) { # 获取指定列范围的列名 cols <- names(dt)[between(names(dt), col_low, col_high)] # 批量匹配+行级判断 dt[, (factor) := as.integer(rowSums(grepl(pattern, .SD), na.rm = TRUE) > 0), .SDcols = cols] return(dt) } # 测试调用 result_dt <- apply_factor_dt(df, 'abc', 'column_c', 'column_e', "^abc") print(result_dt)
核心优化点
- 始终优先使用向量化操作,避免逐行循环(如
rowwise()、apply()逐行) - 利用
grepl可直接处理dataframe/矩阵的特性,批量生成逻辑值 - 用
rowSums快速判断每行是否存在匹配(比逐行any()更高效)
内容的提问来源于stack exchange,提问作者acm_myk
相关产品推荐
相关产品推荐

