如何用separate类函数对DataFrame所有列执行拆分操作?
批量拆分DataFrame中所有
x_y格式列的解决方案 问题场景
现有一个DataFrame,所有列的内容都遵循x_y格式,需要将每一列拆分为两列。单独对某一列使用separate_wider_delim可正常完成拆分,但结合mutate和across批量处理所有列时触发报错,提示data必须是数据框而非字符向量。
测试数据
test <- structure(list(A = c("511686_0.112", "503316_0.105", "476729_0.148", "229348_0.181", "385774_0.178", "209277_0.029", "299921_0.124", "486771_0.123", "524146_0.07", "496030_0.119"), B = c("363323_0.103", "260709_0.105", "361361_0.148", "731426_0.181", "222799_0.178", "140296_0.029", "388191_0.124", "500136_0.123", "487344_0.07", "267303_0.119"), C = c("362981_0.103", "260261_0.105", "360912_0.148", "730423_0.181", "222351_0.178", "139847_0.029", "379717_0.124", "499662_0.123", "486869_0.07", "266907_0.119")), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -10L))
单个列可行代码
test2 <- test %>% separate_wider_delim(A, delim = "_", names_sep = "_")
批量拆分报错代码
test3 <- test %>% mutate(across(everything(), separate_wider_delim, delim = "_", names_sep = "_"))
报错信息
Error in `mutate()`: ℹ In argument: `across(everything(), separate_wider_delim, delim = "_", names_sep = "_")`. Caused by error in `across()`: ! Can't compute column `A`. Caused by error in `fn()`: ! `data` must be a data frame, not a character vector. Run `rlang::last_error()` to see where the error occurred.
报错原因
separate_wider_delim是作用于整个数据框的函数,第一个参数要求传入数据框;而across会将每一列作为单独的字符向量传递给目标函数,导致参数类型不匹配,触发报错。
解决方案
方法一:直接在数据框层面批量拆分(推荐)
无需使用mutate和across,直接调用separate_wider_delim并指定所有列即可:
test3 <- test %>% separate_wider_delim(everything(), delim = "_", names_sep = "_")
everything()表示选中所有列进行拆分names_sep = "_"会让拆分后的新列名格式为原列名_1、原列名_2,例如原列A拆分为A_1和A_2
方法二:结合across的兼容写法(仅作参考)
如果一定要用across,需先将单列向量转换为单列数据框,处理后再展开:
library(tidyr) test3 <- test %>% mutate(across(everything(), ~ { # 将单列向量转为单列数据框 tibble(col = .x) %>% separate_wider_delim(col, delim = "_", names_sep = "_") })) %>% # 展开嵌套的列 unnest_wider(everything())
验证结果
执行推荐方法后,原DataFrame的每一列都会被拆分为两列,例如原列A的"511686_0.112"会被拆分为A_1列的511686和A_2列的0.112。
内容的提问来源于stack exchange,提问作者canderson156
相关产品推荐
相关产品推荐

