You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidyr::separate_wider_regex拆分字符串遇问题,求修正正则

问题:拆分R语言数据框中的字符串列

原始数据

需要处理的R语言数据如下:

id <- c("case1", "case19", "case88", "case77")
vec <- c("One_20 (19)",
         "tWo_20 (290)",
         "Three_38 (399)",
         NA)

df <- data.frame(id, vec)

当前尝试的代码

使用tidyr::separate_wider_regex拆分vec列,但结果不符合预期:

df |> tidyr::separate_wider_regex(vec, 
                                   c(txt = "[A-Za-z]+", num = "\\d+"),
                                   too_few = "align_start")

期望输出

希望得到如下拆分结果:

id      txt num
1  case1   One_20  19
2 case19   tWo_20 290
3 case88 Three_38 399
4 case77     <NA>  NA

修正方案

原正则表达式的问题在于[A-Za-z]+仅匹配纯字母,无法覆盖One_20这类包含下划线、数字的文本部分。调整正则规则,完整匹配文本和括号内的数字:

df |> tidyr::separate_wider_regex(vec, 
                                   c(
                                     txt = "[A-Za-z0-9_]+",  # 匹配字母、数字、下划线组成的文本
                                     "\\s*\\(",              # 匹配括号前的空格和左括号
                                     num = "\\d+",           # 匹配括号内的数字
                                     "\\)"                   # 匹配右括号
                                   ),
                                   too_few = "align_start")

说明

  • txt部分的正则[A-Za-z0-9_]+可以完整捕获One_20、tWo_20这类包含字母、数字、下划线的内容;
  • 补充\\s*\\(和\\)来匹配字符串中的括号结构,确保正确定位括号内的数字作为num列;
  • 保留too_few = "align_start"参数,保证NA值的行能生成对应的NA结果。

内容的提问来源于stack exchange,提问作者JontroPothon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 12:07:20