如何基于data.table按字典指定的位置、名称与类型提取变量
高效从data.table字符串列批量提取指定位置变量并转换类型
示例数据
library(data.table) library(stringi) DT <- data.table(s=c('a191','b292','c393')) dic <- data.table(varname=c('bla','ble','bli'),start=c(1,2,3),end=c(1,2,4),vartype=c('c','i','i')) DTdesired <- data.table(bla=c('a','b','c'),ble=c(1,2,3),bli=c(91,92,93))
高效解决方案
针对2亿条级别的大数据量,使用data.table原地修改操作:=结合向量化字符串提取函数stri_sub,仅遍历变量字典(10-50次),避免低效的逐行处理:
# 批量添加并转换变量 DT[, (dic$varname) := lapply(seq_len(nrow(dic)), function(i) { # 提取指定位置子串 sub_str <- stri_sub(s, start = dic$start[i], end = dic$end[i]) # 根据vartype转换类型 switch(dic$vartype[i], c = sub_str, # 字符型 i = as.integer(sub_str), # 整数型 n = as.numeric(sub_str), # 数值型 sub_str) # 默认返回字符型 })] # 验证结果 all.equal(DT, DTdesired) # [1] TRUE
性能说明
- 原地修改:
:=操作符直接修改原DT,避免复制2亿条数据的巨大内存开销。 - 向量化处理:
stri_sub是向量化函数,一次性处理所有行的字符串提取,远快于逐行循环。 - 低循环次数:仅遍历字典的10-50行,循环成本可忽略。
扩展说明
如果需要支持更多数据类型,只需在switch中添加对应逻辑:
# 示例:添加日期类型转换 dic_extended <- data.table(varname=c('bld'),start=c(1),end=c(4),vartype=c('d')) DT[, (dic_extended$varname) := lapply(seq_len(nrow(dic_extended)), function(i) { sub_str <- stri_sub(s, start = dic_extended$start[i], end = dic_extended$end[i]) switch(dic_extended$vartype[i], d = as.Date(sub_str, format = "%Y"), # 按需定义日期格式 sub_str) })]
内容的提问来源于stack exchange,提问作者LucasMation
相关产品推荐
相关产品推荐

