You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言data.frame指定列拆分求助:separate_wider_delim执行报错

问题描述

我有一个名为table_8[1:2,]的小型data.frame,结构如下:

structure(list(X1 = c("", ""), X2 = c("Wealth", "Ratio (75/25)"), X3 = c("", "Gini"), X4 = c("Non-Land Wealth", "Ratio (75/25) Gini"), X5 = c("Income", "Ratio (75/25)"), X6 = c("", "Gini"), X7 = c("Consumption", "Ratio (75/25) Gini")), row.names = 1:2, class = "data.frame")

需要将X4和X7列的第二行内容各自拆分为两列,使其格式与X2、X3列一致。尝试了以下代码:

table_8_n <- table_8[1:2,] |>
mutate(across(X4:X7, ~str_replace_all(.x, ") G", ")_G")))|>
separate_wider_delim(cols = c(X4,X7), names=LETTERS[1:2], names_sep = "", delim="_", too_few = "align_start")

但出现报错:

Error in `mutate()`:
ℹ In argument: `across(X4:X7, ~str_replace_all(.x, ") G", ")_G"))`.
Caused by error in `across()`:
! Can't compute column `X4`.
Caused by error in `stri_replace_all_regex()`:
! Incorrectly nested parentheses in regex pattern. (U_REGEX_MISMATCHED_PAREN, context=`) G`)

不确定该替换across函数还是用其他方式解决。

解决方案

错误根源

报错是因为str_replace_all的第一个参数是正则表达式,而)属于正则特殊字符,未转义会被当作分组结束符,导致括号匹配错误。另外你不需要用across处理X4到X7的所有列,只需要处理X4和X7即可。

方法一:转义正则特殊字符

把)转义为\\),让正则引擎识别为普通字符,同时缩小across的处理范围:

library(tidyverse)

table_8 <- structure(list(X1 = c("", ""), X2 = c("Wealth", "Ratio (75/25)"), X3 = c("", "Gini"), X4 = c("Non-Land Wealth", "Ratio (75/25) Gini"), X5 = c("Income", "Ratio (75/25)"), X6 = c("", "Gini"), X7 = c("Consumption", "Ratio (75/25) Gini")), row.names = 1:2, class = "data.frame")

table_8_n <- table_8 |>
  mutate(across(c(X4, X7), ~str_replace_all(.x, "\\) G", "\\)_G"))) |>
  separate_wider_delim(
    cols = c(X4, X7), 
    names = c(str_remove(cur_column(), "X"), str_remove(cur_column(), "X")), 
    delim = "_", 
    too_few = "align_start"
  )

这里用cur_column()自动生成对应列名,避免用LETTERS导致命名混乱,同时只处理需要修改的X4和X7列。

方法二:直接处理目标行,规避正则

更直观的方式是定位X4和X7的第二行,直接拆分后插入新列,不需要处理正则转义:

library(tidyverse)

table_8 <- structure(list(X1 = c("", ""), X2 = c("Wealth", "Ratio (75/25)"), X3 = c("", "Gini"), X4 = c("Non-Land Wealth", "Ratio (75/25) Gini"), X5 = c("Income", "Ratio (75/25)"), X6 = c("", "Gini"), X7 = c("Consumption", "Ratio (75/25) Gini")), row.names = 1:2, class = "data.frame")

# 拆分X4列并生成新列
x4_split <- str_split(table_8$X4[2], " ", n=2)[[1]]
table_8 <- table_8 |>
  mutate(X4a = if_else(row_number() == 1, "", x4_split[1]),
         X4b = if_else(row_number() == 1, "Gini", x4_split[2])) |>
  select(-X4)

# 拆分X7列并生成新列
x7_split <- str_split(table_8$X7[2], " ", n=2)[[1]]
table_8 <- table_8 |>
  mutate(X7a = if_else(row_number() == 1, "", x7_split[1]),
         X7b = if_else(row_number() == 1, "Gini", x7_split[2])) |>
  select(-X7)

# 调整列顺序,对齐原表格格式
table_8 <- table_8 |>
  select(X1, X2, X3, X4a, X4b, X5, X6, X7a, X7b)

这种方法直接针对需要修改的内容操作,逻辑更清晰,也不会遇到正则特殊字符的问题。

内容的提问来源于stack exchange,提问作者hks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 20:07:50