You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R的Tidymodels Recipe中如何正确按行计算正预测变量数并缩放?

问题

需要在tidymodels的Recipe中实现:将特定一组预测变量,除以该行内该组变量中大于0的数量。当前实现中变量被错误除以该组变量的总数,而非行内符合条件的数量。

错误原因

原Recipe中的sum(all_of(vars) > 0)是全局求和,计算的是整个数据集里该组变量大于0的总个数,而非逐行统计。另外,生成的pos_count会被纳入模型变量,干扰拟合结果。

解决方案
  1. 用rowSums()结合across()逐行统计目标变量组中大于0的数量;
  2. 完成变量除法操作后,用step_rm()移除pos_count,避免其进入模型;
  3. 保持across()与all_of(vars)的组合,精准定位目标变量组。
修正后的完整代码
df <- data.frame(matrix(c(16, 8, 4, 2, 32, 16, 8, 4, 0, 32, 16, 8, 0, 0, 32, 16, 0, 0, 0, 32), 4, 5))
vars <- names(df)[-1]

lm_model <- 
  linear_reg(penalty = 0) %>%  
  set_engine("glmnet", lower.limits = rep(0, 5), upper.limits = rep(1, 5), intercept = FALSE)

# 修正后的Recipe
lm_recipe <- 
  recipe(X1 ~ X2 + X3 + X4 + X5, data = df) %>% 
  # 逐行统计目标变量组中大于0的数量
  step_mutate(pos_count = rowSums(across(all_of(vars), ~. > 0))) %>%
  # 对目标变量组执行除法操作
  step_mutate(across(all_of(vars), ~ . / pos_count)) %>%
  # 移除pos_count,避免进入模型
  step_rm(pos_count)

lm_wflow <- 
  workflow() %>% 
  add_model(lm_model) %>%
  add_recipe(lm_recipe)

lm_fit <- fit(lm_wflow, df)
lm_fit %>% tidy()
修正后的输出
term        estimate penalty
1 (Intercept)   0            0
2 X2            0.492        0
3 X3            0.240        0
4 X4            0.112        0
5 X5            0.0256       0

该结果与手动预处理数据后的拟合结果完全一致。

内容的提问来源于stack exchange,提问作者ThePhil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 05:27:25