从dplyr转用data.table:自定义函数调用报错及修改咨询
解决data.table自定义函数中字符串列名引用的报错问题
我此前使用dplyr,刚接触data.table,编写了如下自定义函数tabs:
tabs <- function(dt, x) { tab2 <- dt[!is.na(x), ][, .(Freq = sum(nwgt0)), by = .(inc_cat, year, x)][, Prop := Freq / sum(Freq), by= .(inc_cat, year)][order(inc_cat, year)][x == 1 & !is.na(inc_cat), ] %>% ggplot(., aes(x= year, y = Prop, color = factor(inc_cat, levels = c(1,2,3,4),labels = c("0% to 100% FPL", "101-138% FPL", "139-200% FPL", ">200% FPL")))) + labs(color = "Income Categories") + geom_line() + theme_minimal() + ylab("Weighted proportion") + theme( panel.border = element_blank(), panel.grid.major = element_blank(), panel.grid.minor = element_blank(), ) return(tab2) }
尝试通过result <- hints_dt[ , tabs(.SD, x='internet_use')]调用时出现报错:
Error in `[.data.table`(dt[!is.na(x), ], , .(Freq = sum(nwgt0)), by = .(inc_cat, : The items in the 'by' or 'keyby' list are length(s) (22344,22344,1). Each must be length 22344; the same length as there are rows in x (after subsetting if i is provided).
附测试用reprex(基于NHANES数据集):
tabs <- function(dt, x) { tab2 <- dt[!is.na(x), ][, .(Freq = sum(WTMEC2YR)), by = .(race, agecat, x)][, Prop := Freq / sum(Freq), by= .(race, agecat)][order(race, agecat)][x == 1 & !is.na(race), ] %>% ggplot(., aes(x= year, y = Prop, color = factor(race, levels = c(1,2,3,4),labels = c("hispanic", "white", "black", "other")))) + labs(color = "Race") + geom_line() + theme_minimal() + ylab("Weighted proportion") + theme( panel.border = element_blank(), panel.grid.major = element_blank(), panel.grid.minor = element_blank(), ) return(tab2) }
调用result <- nhanes[ , tabs(.SD, x="RIAGENDR")]时可复现报错:
Error in `[.data.table`(dt[!is.na(get(x)), ], , .(Freq = sum(WTMEC2YR)), : The items in the 'by' or 'keyby' list are length(s) (8591,8591,1). Each must be length 8591; the same length as there are rows in x (after subsetting if i is provided).
错误原因
报错核心是直接用字符串变量x引用列名不符合data.table语法规则:
dt[!is.na(x)]中,x是传入的字符串(如"internet_use"),并非列对象,!is.na(x)会返回长度为1的逻辑值,导致子集筛选逻辑错误;by = .(inc_cat, year, x)里的x同样是字符串,长度为1,和inc_cat、year的行长度不匹配,触发分组长度不一致的报错。
修改方案
用get(x)动态引用列名,确保data.table能识别到目标列;同时可结合.SDcols优化列传递(可选,减少不必要的列加载)。
修改后的原始函数
tabs <- function(dt, x) { tab2 <- dt[!is.na(get(x)), ] %>% .[, .(Freq = sum(nwgt0)), by = .(inc_cat, year, get(x))] %>% .[, Prop := Freq / sum(Freq), by = .(inc_cat, year)] %>% .[order(inc_cat, year)] %>% .[get(x) == 1 & !is.na(inc_cat), ] %>% ggplot(aes(x = year, y = Prop, color = factor(inc_cat, levels = c(1,2,3,4), labels = c("0% to 100% FPL", "101-138% FPL", "139-200% FPL", ">200% FPL")))) + labs(color = "Income Categories") + geom_line() + theme_minimal() + ylab("Weighted proportion") + theme( panel.border = element_blank(), panel.grid.major = element_blank(), panel.grid.minor = element_blank() ) return(tab2) }
修改后的测试用函数(NHANES数据集)
tabs <- function(dt, x) { tab2 <- dt[!is.na(get(x)), ] %>% .[, .(Freq = sum(WTMEC2YR)), by = .(race, agecat, get(x))] %>% .[, Prop := Freq / sum(Freq), by = .(race, agecat)] %>% .[order(race, agecat)] %>% .[get(x) == 1 & !is.na(race), ] %>% ggplot(aes(x = year, y = Prop, color = factor(race, levels = c(1,2,3,4), labels = c("hispanic", "white", "black", "other")))) + labs(color = "Race") + geom_line() + theme_minimal() + ylab("Weighted proportion") + theme( panel.border = element_blank(), panel.grid.major = element_blank(), panel.grid.minor = element_blank() ) return(tab2) }
调用方式优化(可选)
指定.SDcols只传递需要的列,提升效率:
# 原始数据调用 result <- hints_dt[, tabs(.SD, x='internet_use'), .SDcols = c('inc_cat', 'year', 'nwgt0', 'internet_use')] # NHANES测试数据调用 result <- nhanes[, tabs(.SD, x="RIAGENDR"), .SDcols = c('race', 'agecat', 'year', 'WTMEC2YR', 'RIAGENDR')]
内容的提问来源于stack exchange,提问作者Felippe Marcondes
相关产品推荐
相关产品推荐

