You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则匹配字符串中最后一个逗号以分割数据列?

解决方案

你的问题出在separate默认会匹配所有符合sep规则的分隔符,导致带逗号的公司名被错误拆分。要只分割最后一个逗号,可以用以下两种方法:

方法1:使用正则表达式匹配最后一个逗号

利用正则正向预查,匹配后面没有其他逗号的, :

library(tidyverse)

df <- tibble(
  name = c("John", "James"), 
  company_num = c("Apple, Inc, 1000",
                  "Microsoft, 1200")
)

df %>% 
  separate(col = company_num, 
           into = c("company", "num"), 
           sep = ", (?=[^,]+$)",  # 匹配最后一个", "
           convert = TRUE)  # 自动将num转为数值型

方法2:使用separate_wider_delim(tidyverse 1.0.0+支持)

这个函数专门支持按最后一个分隔符拆分,逻辑更直观:

df %>% 
  separate_wider_delim(
    cols = company_num,
    delim = ", ",
    names = c("company", "num"),
    too_few = "align_start",  # 适配只有一个分隔符的条目
    convert = TRUE
  )

两种方法都能得到目标输出:

# A tibble: 2 x 3
  name  company      num
  <chr> <chr>      <dbl>
1 John  Apple, Inc  1000
2 James Microsoft   1200

内容的提问来源于stack exchange,提问作者HoelR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 13:05:24