You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

tidyr包中separate_longer_delim报错:无法循环输入,如何修复?

问题:使用separate_longer_delim时出现回收错误

当运行以下R代码调用separate_longer_delim函数时,触发报错:

In row 2, can't recycle input of size 4 to size 6.

原始代码如下:

library(tidyverse)
ori_df <- data.frame(
  Cat_A = c("BDW","A_B_W_F"),
  Cat_B = c("AWS","OVS"),
  Cat_C = c("CATA","CATB"),
  Cat_D = c("ABCDE","PBCD"),
  Weight = c("1kg_2kg_3kg_4kg","1kg_2kg_3kg_5kg_10kg_20kg"),
  Country = c("澳洲_德国_法国_加拿大_美国_日本_西班牙_意大利_英国",
              "巴西_德国_俄罗斯_法国_加拿大_美国_日本_西班牙_意大利_英国"))

ori_df %>% separate_longer_delim(c('Cat_A','Weight','Country'),delim="_")

错误原因

separate_longer_delim要求同一行内,所有指定拆分的列,拆分后的元素数量必须完全一致,这样才能按位置配对展开成多行。但你的数据第2行中:

  • Cat_A拆分后有4个元素(A、B、W、F)
  • Weight拆分后有6个元素
  • Country拆分后有10个元素
    元素数量不匹配,函数无法完成回收配对,因此抛出错误。

修复方案

方案1:保留所有元素,短列表补NA

如果需要保留所有拆分后的元素,对长度不足的列用NA填充,可以结合rowwise()、str_split和unnest_longer实现:

ori_df %>%
  rowwise() %>%
  mutate(
    Cat_A = list(str_split(Cat_A, "_")[[1]]),
    Weight = list(str_split(Weight, "_")[[1]]),
    Country = list(str_split(Country, "_")[[1]])
  ) %>%
  unnest_longer(c(Cat_A, Weight, Country), keep_empty = TRUE)

方案2:仅保留元素数量匹配的行

如果业务逻辑要求同一行的拆分列元素数量必须一致,先检查每行的拆分长度,再过滤或调整数据:

  1. 查看每行各列的拆分长度:
ori_df %>%
  rowwise() %>%
  mutate(
    len_CatA = length(str_split(Cat_A, "_")[[1]]),
    len_Weight = length(str_split(Weight, "_")[[1]]),
    len_Country = length(str_split(Country, "_")[[1]])
  ) %>%
  ungroup()
  1. 过滤长度匹配的行(示例,需根据实际业务调整):
ori_df %>%
  rowwise() %>%
  filter(
    length(str_split(Cat_A, "_")[[1]]) == length(str_split(Weight, "_")[[1]]) &
    length(str_split(Weight, "_")[[1]]) == length(str_split(Country, "_")[[1]])
  ) %>%
  ungroup() %>%
  separate_longer_delim(c('Cat_A','Weight','Country'), delim="_")

方案3:循环短列表对齐长列表

如果希望让短列表循环重复,以匹配最长列表的长度,可以用rep调整后再展开:

ori_df %>%
  rowwise() %>%
  mutate(
    cat_a_list = str_split(Cat_A, "_")[[1]],
    weight_list = str_split(Weight, "_")[[1]],
    country_list = str_split(Country, "_")[[1]],
    max_len = max(length(cat_a_list), length(weight_list), length(country_list)),
    Cat_A = list(rep(cat_a_list, length.out = max_len)),
    Weight = list(rep(weight_list, length.out = max_len)),
    Country = list(rep(country_list, length.out = max_len))
  ) %>%
  select(-cat_a_list, -weight_list, -country_list, -max_len) %>%
  unnest_longer(c(Cat_A, Weight, Country))

内容的提问来源于stack exchange,提问作者anderwyang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 04:57:10