如何在多列大data.table中筛选含≥2值的行及对应符合条件的列
解决方案
针对大尺寸data.table的场景,更推荐使用原生向量化操作实现需求,避免apply按行遍历带来的性能损耗,完整实现逻辑如下:
分步实现代码
- 加载依赖&构造示例数据
library(data.table) DT <- as.data.table( rbind( diag(5)+1, diag(5), rep(0,5)) ) DT[1,1] = 1
- 第一步:筛选至少存在一个值≥2的行
# 用Reduce做列向布尔运算,性能远高于按行apply row_cond <- DT[, Reduce(`|`, lapply(.SD, `>=`, 2))] dt_row_filtered <- DT[row_cond]
- 第二步:筛选包含至少一个值≥2的列
dt_final <- dt_row_filtered[, .SD, .SDcols = sapply(.SD, function(x) any(x >= 2))]
结果验证
运行dt_final即可得到预期输出:
V2 V3 V4 V5 1: 2 1 1 1 2: 1 2 1 1 3: 1 1 2 1 4: 1 1 1 2
兼容原有写法的方案
如果你已经在使用@IRTFM提供的行筛选代码,直接补充.SDcols参数筛选列即可:
dt_row_filtered <- DT[DT[,apply(.SD, 1, function(x) any(x>=2))]] dt_final <- dt_row_filtered[, .SD, .SDcols = sapply(dt_row_filtered, function(x) any(x >= 2))]
内容的提问来源于stack exchange,提问作者moth
相关产品推荐
相关产品推荐

