为何同一列设两个条件时dplyr的case_when无法正常工作?
问题描述
在dplyr中使用case_when函数,根据abif_type列的值从不同数据表匹配对应数据。单独使用任一条件均可正常运行,但同时添加两个条件后,触发错误:
'names' attribute [1] must be the same length as the vector [0]
以下是复现错误的完整代码:
library(lubridate) directory <- structure(list(data_label = c("BufT.1", "CTID.1", "CTNM.1", "CTOw.1", "CTTL.1", "CpEP.1", "DATA.1", "DATA.2", "DATA.3", "DATA.4", "DATA.5", "DATA.6", "DATA.7", "DATA.8", "DATA.105", "DCHT.1", "DSam.1", "DySN.1", "Dye#.1", "DyeB.1", "DyeB.2", "DyeB.3", "DyeB.4", "DyeB.5", "DyeN.1", "DyeN.2", "DyeN.3", "DyeN.4", "DyeN.5", "DyeW.1", "DyeW.2", "DyeW.3", "DyeW.4", "DyeW.5", "EPVt.1", "EVNT.1", "EVNT.2", "EVNT.3", "EVNT.4", "GTyp.1", "HCFG.1", "HCFG.2", "HCFG.3", "HCFG.4", "InSc.1", "InVt.1", "LANE.1", "LIMS.1", "LNTD.1", "LsrP.1", "MCHN.1", "MODF.1", "MODL.1", "NAVG.1", "NLNE.1", "OfSc.1", "PSZE.1", "PTYP.1", "PXLB.1", "RGNm.1", "RGOw.1", "RMXV.1", "RMdN.1", "RMdV.1", "RMdX.1", "RPrN.1", "RPrV.1", "RUND.1", "RUND.2", "RUND.3", "RUND.4", "RUNT.1", "RUNT.2", "RUNT.3", "RUNT.4", "Rate.1", "RunN.1", "SCAN.1", "SMED.1", "SMLt.1", "SVER.1", "SVER.3", "SVER.4", "Satd.1", "Scal.1", "Scan.1", "SpNm.1", "TUBE.1", "Tmpr.1", "User.1"), abif_type = c("short array", "cString", "cString", "cString", "pString", "byte", "short array", "short array", "short array", "short array", "short array", "short array", "short array", "short array", "short array", "short", "short", "pString", "short", "char", "char", "char", "char", "char", "pString", "pString", "pString", "pString", "pString", "short", "short", "short", "short", "short", "long", "pString", "pString", "pString", "pString", "pString", "cString", "cString", "cString", "cString", "long", "long", "short", "pString", "short", "long", "pString", "pString", "char[4]", "short", "short", "long array", "long", "cString", "long", "cString", "cString", "cString", "cString", "cString", "char array", "cString", "cString", "date", "date", "date", "date", "time", "time", "time", "time", "user", "cString", "long", "pString", "pString", "pString", "pString", "pString", "long array", "float", "short", "pString", "pString", "long", "pString")), row.names = c(NA, -90L), class = "data.frame") data.date <- structure(list(index = c("RUND.1", "RUND.2", "RUND.3", "RUND.4" ), date = structure(c(19234, 19234, 19234, 19234), class = "Date")), row.names = c(NA, -4L), class = c("tbl_df", "tbl", "data.frame")) data.time <- structure(list(index = c("RUNT.1", "RUNT.2", "RUNT.3", "RUNT.4" ), time = new("Period", .Data = c(56, 47, 8, 48), year = c(0, 0, 0, 0), month = c(0, 0, 0, 0), day = c(0, 0, 0, 0), hour = c(17, 19, 18, 19), minute = c(48, 0, 19, 0)), hsecond = c(0L, 0L, 0L, 0L)), row.names = c(NA, -4L), class = c("tbl_df", "tbl", "data.frame" )) library(dplyr) results <- directory %>% mutate(value = case_when( abif_type == "date" ~ data.date$date[match(data_label, data.date$index)], abif_type == "time" ~ data.time$time[match(data_label, data.time$index)]))
问题原因与解决方案
核心原因
case_when有两个严格要求:
- 所有分支返回的结果数据类型必须一致;
- 每个分支返回的结果长度必须与输入数据的行数匹配,且不能出现属性长度不匹配的情况。
在你的代码中,data.date$date是Date类型,data.time$time是lubridate::Period类型,二者类型不兼容;同时,当某行不满足当前分支条件时,match返回的NA在结合不同类型的向量时,会触发内部的属性长度检查错误。
解决方案
方法1:统一数据类型并设置默认值
将两种类型统一转为字符型(或其他通用类型),同时为不满足条件的行明确设置默认值:
results <- directory %>% mutate(value = case_when( abif_type == "date" ~ as.character(data.date$date[match(data_label, data.date$index)]), abif_type == "time" ~ as.character(data.time$time[match(data_label, data.time$index)]), TRUE ~ NA_character_ # 覆盖所有其他情况 ))
方法2:用左连接替代case_when(更稳妥)
先通过左连接将两个数据表的内容合并到主表中,再用case_when选择对应的值:
results <- directory %>% left_join(data.date, by = c("data_label" = "index")) %>% left_join(data.time, by = c("data_label" = "index")) %>% mutate(value = case_when( abif_type == "date" ~ as.character(date), abif_type == "time" ~ as.character(time), TRUE ~ NA_character_ )) %>% select(-date, -time) # 移除临时生成的列
方法3:保留原始类型(用列表存储)
如果需要保留Date和Period的原始类型,可以将结果存储为列表,再展开:
results <- directory %>% mutate(value = case_when( abif_type == "date" ~ list(data.date$date[match(data_label, data.date$index)]), abif_type == "time" ~ list(data.time$time[match(data_label, data.time$index)]), TRUE ~ list(NA) )) %>% tidyr::unnest(value)
内容的提问来源于stack exchange,提问作者Mike
相关产品推荐
相关产品推荐

