如何使用Tidymodels Recipe按行逻辑计算衍生列并筛选数据
实现代码
# 加载所需工具包 library(tidymodels) # 构造你的示例数据 df <- tibble( ID = 1:5, Col1 = c("blue", "orange", "red", "yellow", "green"), RespID = c("729Ad", "295gS", "729Ad", "592Jd", "937sa"), Col3 = c(3.2, 6.5, 8.4, 2.9, 3.5), Col4 = c("A", "A", "B", "A", "B") ) # 构建完整处理流程 rec <- recipe(~ ., data = df) %>% # 按RespID分组计算Col5 step_mutate( Col5 = as.integer(Col4 == "A" & any(Col4 == "B")), .by = RespID ) %>% # 过滤仅保留Col4为A的行 step_filter(Col4 == "A") # 执行流程得到最终结果 df_final <- rec %>% prep(training = df) %>% bake(new_data = NULL)
生成删除B行前的中间表
如果需要验证你给出的预期中间结果,只需去掉step_filter步骤即可:
rec_mid <- recipe(~ ., data = df) %>% step_mutate( Col5 = as.integer(Col4 == "A" & any(Col4 == "B")), .by = RespID ) df_mid <- rec_mid %>% prep(training = df) %>% bake(new_data = NULL)
逻辑说明
step_mutate中的.by = RespID表示按RespID分组执行计算,any(Col4 == "B")会判断当前分组内是否存在Col4为B的记录- 两个判断条件
当前行Col4为A+同RespID下存在B行同时满足时返回TRUE,转整数后为1,其余所有情况返回0,完全匹配你要求的Col5计算规则 step_filter直接过滤掉所有Col4为B的行,仅保留A类行
内容的提问来源于stack exchange,提问作者piper180
相关产品推荐
相关产品推荐

