tidyr包中separate_longer_delim报错:无法循环输入,如何修复?
问题:使用
separate_longer_delim时出现回收错误 当运行以下R代码调用separate_longer_delim函数时,触发报错:
In row 2, can't recycle input of size 4 to size 6.
原始代码如下:
library(tidyverse) ori_df <- data.frame( Cat_A = c("BDW","A_B_W_F"), Cat_B = c("AWS","OVS"), Cat_C = c("CATA","CATB"), Cat_D = c("ABCDE","PBCD"), Weight = c("1kg_2kg_3kg_4kg","1kg_2kg_3kg_5kg_10kg_20kg"), Country = c("澳洲_德国_法国_加拿大_美国_日本_西班牙_意大利_英国", "巴西_德国_俄罗斯_法国_加拿大_美国_日本_西班牙_意大利_英国")) ori_df %>% separate_longer_delim(c('Cat_A','Weight','Country'),delim="_")
错误原因
separate_longer_delim要求同一行内,所有指定拆分的列,拆分后的元素数量必须完全一致,这样才能按位置配对展开成多行。但你的数据第2行中:
Cat_A拆分后有4个元素(A、B、W、F)Weight拆分后有6个元素Country拆分后有10个元素
元素数量不匹配,函数无法完成回收配对,因此抛出错误。
修复方案
方案1:保留所有元素,短列表补NA
如果需要保留所有拆分后的元素,对长度不足的列用NA填充,可以结合rowwise()、str_split和unnest_longer实现:
ori_df %>% rowwise() %>% mutate( Cat_A = list(str_split(Cat_A, "_")[[1]]), Weight = list(str_split(Weight, "_")[[1]]), Country = list(str_split(Country, "_")[[1]]) ) %>% unnest_longer(c(Cat_A, Weight, Country), keep_empty = TRUE)
方案2:仅保留元素数量匹配的行
如果业务逻辑要求同一行的拆分列元素数量必须一致,先检查每行的拆分长度,再过滤或调整数据:
- 查看每行各列的拆分长度:
ori_df %>% rowwise() %>% mutate( len_CatA = length(str_split(Cat_A, "_")[[1]]), len_Weight = length(str_split(Weight, "_")[[1]]), len_Country = length(str_split(Country, "_")[[1]]) ) %>% ungroup()
- 过滤长度匹配的行(示例,需根据实际业务调整):
ori_df %>% rowwise() %>% filter( length(str_split(Cat_A, "_")[[1]]) == length(str_split(Weight, "_")[[1]]) & length(str_split(Weight, "_")[[1]]) == length(str_split(Country, "_")[[1]]) ) %>% ungroup() %>% separate_longer_delim(c('Cat_A','Weight','Country'), delim="_")
方案3:循环短列表对齐长列表
如果希望让短列表循环重复,以匹配最长列表的长度,可以用rep调整后再展开:
ori_df %>% rowwise() %>% mutate( cat_a_list = str_split(Cat_A, "_")[[1]], weight_list = str_split(Weight, "_")[[1]], country_list = str_split(Country, "_")[[1]], max_len = max(length(cat_a_list), length(weight_list), length(country_list)), Cat_A = list(rep(cat_a_list, length.out = max_len)), Weight = list(rep(weight_list, length.out = max_len)), Country = list(rep(country_list, length.out = max_len)) ) %>% select(-cat_a_list, -weight_list, -country_list, -max_len) %>% unnest_longer(c(Cat_A, Weight, Country))
内容的提问来源于stack exchange,提问作者anderwyang
相关产品推荐
相关产品推荐

