R语言data.frame指定列拆分求助:separate_wider_delim执行报错
问题描述
我有一个名为table_8[1:2,]的小型data.frame,结构如下:
structure(list(X1 = c("", ""), X2 = c("Wealth", "Ratio (75/25)"), X3 = c("", "Gini"), X4 = c("Non-Land Wealth", "Ratio (75/25) Gini"), X5 = c("Income", "Ratio (75/25)"), X6 = c("", "Gini"), X7 = c("Consumption", "Ratio (75/25) Gini")), row.names = 1:2, class = "data.frame")
需要将X4和X7列的第二行内容各自拆分为两列,使其格式与X2、X3列一致。尝试了以下代码:
table_8_n <- table_8[1:2,] |> mutate(across(X4:X7, ~str_replace_all(.x, ") G", ")_G")))|> separate_wider_delim(cols = c(X4,X7), names=LETTERS[1:2], names_sep = "", delim="_", too_few = "align_start")
但出现报错:
Error in `mutate()`: ℹ In argument: `across(X4:X7, ~str_replace_all(.x, ") G", ")_G"))`. Caused by error in `across()`: ! Can't compute column `X4`. Caused by error in `stri_replace_all_regex()`: ! Incorrectly nested parentheses in regex pattern. (U_REGEX_MISMATCHED_PAREN, context=`) G`)
不确定该替换across函数还是用其他方式解决。
解决方案
错误根源
报错是因为str_replace_all的第一个参数是正则表达式,而)属于正则特殊字符,未转义会被当作分组结束符,导致括号匹配错误。另外你不需要用across处理X4到X7的所有列,只需要处理X4和X7即可。
方法一:转义正则特殊字符
把)转义为\\),让正则引擎识别为普通字符,同时缩小across的处理范围:
library(tidyverse) table_8 <- structure(list(X1 = c("", ""), X2 = c("Wealth", "Ratio (75/25)"), X3 = c("", "Gini"), X4 = c("Non-Land Wealth", "Ratio (75/25) Gini"), X5 = c("Income", "Ratio (75/25)"), X6 = c("", "Gini"), X7 = c("Consumption", "Ratio (75/25) Gini")), row.names = 1:2, class = "data.frame") table_8_n <- table_8 |> mutate(across(c(X4, X7), ~str_replace_all(.x, "\\) G", "\\)_G"))) |> separate_wider_delim( cols = c(X4, X7), names = c(str_remove(cur_column(), "X"), str_remove(cur_column(), "X")), delim = "_", too_few = "align_start" )
这里用cur_column()自动生成对应列名,避免用LETTERS导致命名混乱,同时只处理需要修改的X4和X7列。
方法二:直接处理目标行,规避正则
更直观的方式是定位X4和X7的第二行,直接拆分后插入新列,不需要处理正则转义:
library(tidyverse) table_8 <- structure(list(X1 = c("", ""), X2 = c("Wealth", "Ratio (75/25)"), X3 = c("", "Gini"), X4 = c("Non-Land Wealth", "Ratio (75/25) Gini"), X5 = c("Income", "Ratio (75/25)"), X6 = c("", "Gini"), X7 = c("Consumption", "Ratio (75/25) Gini")), row.names = 1:2, class = "data.frame") # 拆分X4列并生成新列 x4_split <- str_split(table_8$X4[2], " ", n=2)[[1]] table_8 <- table_8 |> mutate(X4a = if_else(row_number() == 1, "", x4_split[1]), X4b = if_else(row_number() == 1, "Gini", x4_split[2])) |> select(-X4) # 拆分X7列并生成新列 x7_split <- str_split(table_8$X7[2], " ", n=2)[[1]] table_8 <- table_8 |> mutate(X7a = if_else(row_number() == 1, "", x7_split[1]), X7b = if_else(row_number() == 1, "Gini", x7_split[2])) |> select(-X7) # 调整列顺序,对齐原表格格式 table_8 <- table_8 |> select(X1, X2, X3, X4a, X4b, X5, X6, X7a, X7b)
这种方法直接针对需要修改的内容操作,逻辑更清晰,也不会遇到正则特殊字符的问题。
内容的提问来源于stack exchange,提问作者hks
相关产品推荐
相关产品推荐

