You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Tidymodels Recipe按行逻辑计算衍生列并筛选数据

实现代码

# 加载所需工具包
library(tidymodels)

# 构造你的示例数据
df <- tibble(
  ID = 1:5,
  Col1 = c("blue", "orange", "red", "yellow", "green"),
  RespID = c("729Ad", "295gS", "729Ad", "592Jd", "937sa"),
  Col3 = c(3.2, 6.5, 8.4, 2.9, 3.5),
  Col4 = c("A", "A", "B", "A", "B")
)

# 构建完整处理流程
rec <- recipe(~ ., data = df) %>%
  # 按RespID分组计算Col5
  step_mutate(
    Col5 = as.integer(Col4 == "A" & any(Col4 == "B")),
    .by = RespID
  ) %>%
  # 过滤仅保留Col4为A的行
  step_filter(Col4 == "A")

# 执行流程得到最终结果
df_final <- rec %>%
  prep(training = df) %>%
  bake(new_data = NULL)

生成删除B行前的中间表

如果需要验证你给出的预期中间结果,只需去掉step_filter步骤即可:

rec_mid <- recipe(~ ., data = df) %>%
  step_mutate(
    Col5 = as.integer(Col4 == "A" & any(Col4 == "B")),
    .by = RespID
  )

df_mid <- rec_mid %>%
  prep(training = df) %>%
  bake(new_data = NULL)

逻辑说明

  • step_mutate中的.by = RespID表示按RespID分组执行计算,any(Col4 == "B")会判断当前分组内是否存在Col4为B的记录
  • 两个判断条件当前行Col4为A+同RespID下存在B行同时满足时返回TRUE,转整数后为1,其余所有情况返回0,完全匹配你要求的Col5计算规则
  • step_filter直接过滤掉所有Col4为B的行,仅保留A类行

内容的提问来源于stack exchange,提问作者piper180

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 13:09:02