在R的Tidymodels Recipe中如何正确按行计算正预测变量数并缩放?
问题
需要在tidymodels的Recipe中实现:将特定一组预测变量,除以该行内该组变量中大于0的数量。当前实现中变量被错误除以该组变量的总数,而非行内符合条件的数量。
错误原因
原Recipe中的sum(all_of(vars) > 0)是全局求和,计算的是整个数据集里该组变量大于0的总个数,而非逐行统计。另外,生成的pos_count会被纳入模型变量,干扰拟合结果。
解决方案
- 用
rowSums()结合across()逐行统计目标变量组中大于0的数量; - 完成变量除法操作后,用
step_rm()移除pos_count,避免其进入模型; - 保持
across()与all_of(vars)的组合,精准定位目标变量组。
修正后的完整代码
df <- data.frame(matrix(c(16, 8, 4, 2, 32, 16, 8, 4, 0, 32, 16, 8, 0, 0, 32, 16, 0, 0, 0, 32), 4, 5)) vars <- names(df)[-1] lm_model <- linear_reg(penalty = 0) %>% set_engine("glmnet", lower.limits = rep(0, 5), upper.limits = rep(1, 5), intercept = FALSE) # 修正后的Recipe lm_recipe <- recipe(X1 ~ X2 + X3 + X4 + X5, data = df) %>% # 逐行统计目标变量组中大于0的数量 step_mutate(pos_count = rowSums(across(all_of(vars), ~. > 0))) %>% # 对目标变量组执行除法操作 step_mutate(across(all_of(vars), ~ . / pos_count)) %>% # 移除pos_count,避免进入模型 step_rm(pos_count) lm_wflow <- workflow() %>% add_model(lm_model) %>% add_recipe(lm_recipe) lm_fit <- fit(lm_wflow, df) lm_fit %>% tidy()
修正后的输出
term estimate penalty 1 (Intercept) 0 0 2 X2 0.492 0 3 X3 0.240 0 4 X4 0.112 0 5 X5 0.0256 0
该结果与手动预处理数据后的拟合结果完全一致。
内容的提问来源于stack exchange,提问作者ThePhil
相关产品推荐
相关产品推荐

