在R语言分组tibble中使用if_else()的错误排查与解决咨询
问题:为卫星工厂分配对应主工厂ID
需要基于数据集的年份year、客户IDcustomer、工厂IDplant_code,以及主/卫星工厂标识(is_main/is_satellite),为每个客户-年份组合下的卫星工厂分配主工厂ID(若该组存在主工厂),否则保留原工厂ID,新字段命名为plant_code_new。
示例数据集
library(tidyverse) df <- tibble( year = c(2021, 2021, 2022, 2022, 2021, 2022, 2021, 2022, 2021, 2021, 2022, 2022, 2021, 2022, 2021, 2022), customer = c("A", "A", "A", "A", "B", "B", "C", "C", "D", "D", "D", "D", "E", "E", "E", "E"), plant_code = c(1, 2, 1, 2, 3, 3, 4, 4, 5, 6, 5, 6, 7, 7, 8, 8), is_main = c(1, 0, 1, 0, 1, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1), is_satellite = c(0, 1, 0, 1, 0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0))
尝试的代码及错误
尝试代码
df_adj <- df |> group_by(year, customer) |> mutate(grp_size = (n() - sum(is_satellite))) |> mutate(plant_code_new = if_else(condition = rep((n() - sum(is_satellite)) == 1, times = n()), true = plant_code[is_main == 1], false = plant_code, size = n()))
错误信息
Error in
mutate():
In argument:plant_code_new = if_else(...).
In group 3:year = 2021,customer = "D".
Caused by error inif_else():
!truemust have size 2, not size 0.
问题原因
客户D 2021年的分组中,所有工厂都是卫星工厂(is_main全为0),导致plant_code[is_main == 1]返回长度为0的向量。尽管该组的逻辑测试结果为FALSE,但if_else会先校验所有参数的长度是否匹配分组行数(此处为2),无论条件是否触发对应分支,参数长度不匹配就会报错。
解决方案
换一种思路:先提取每个客户-年份组的主工厂ID,再通过条件判断替换卫星工厂的ID,避免参数长度不匹配问题:
df_adj <- df |> group_by(year, customer) |> # 提取当前组的主工厂ID,无主工厂则返回NA mutate(main_plant = first(plant_code[is_main == 1])) |> # 卫星工厂替换为主工厂ID,其余情况保留原ID mutate(plant_code_new = case_when( is_satellite == 1 & !is.na(main_plant) ~ main_plant, TRUE ~ plant_code )) |> ungroup()
结果说明
运行后各分组的plant_code_new符合预期:
- 客户A的卫星工厂(
plant_code=2)会被替换为主工厂ID 1; - 客户B、C、D、E的工厂要么是主工厂,要么无对应主工厂,因此保留原
plant_code。
内容的提问来源于stack exchange,提问作者Hugo
相关产品推荐
相关产品推荐

