You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R语言将数据框指定列按最后下划线拆分为两列?

解决方案

这里提供几种适配你的tibble数据结构的实现方法:

方法1:使用tidyr包(tidyverse生态)

利用separate_wider_regex可以精准定位最后一个下划线的位置完成拆分,无最后下划线的行自动填充NA:

library(tidyverse)

df <- tibble(col1 = c("2397_A_merge", "3779_A_merge", "4958_BV"))

df_new <- df %>%
  separate_wider_regex(
    col1,
    patterns = list(
      col1 = "^.*(?=_)",
      col2 = "(?<=_).*$"
    ),
    too_few = "align_start"
  )

df_new

运行结果:

# A tibble: 3 × 2
  col1    col2 
  <chr>   <chr>
1 2397_A  merge
2 3779_A  merge
3 4958_BV NA   

正则说明:

  • ^.*(?=_):匹配从开头到最后一个下划线之前的所有字符(正向预查确保后续是下划线)
  • (?<=_).*$:匹配最后一个下划线之后到结尾的所有字符(反向预查确保前面是下划线)
  • too_few = "align_start":未匹配到第二部分时,将第二列设为NA

也可以用更简洁的separate函数,通过正则指定拆分位置为最后一个下划线:

df %>%
  separate(
    col1,
    into = c("col1", "col2"),
    sep = "(?<=.)_(?=[^_]+$)",
    fill = "right"
  )

sep = "(?<=.)_(?=[^_]+$)"仅匹配后续无其他下划线的那个下划线,fill = "right"确保右侧无内容时填充NA。

方法2:使用Base R

无需额外加载包,通过基础字符串处理函数实现:

df <- tibble(col1 = c("2397_A_merge", "3779_A_merge", "4958_BV"))

# 提取最后下划线后的内容作为col2
df$col2 <- sub("^.*_", "", df$col1)
# 保留最后下划线前的内容作为新col1
df$col1 <- sub("_(.*)$", "", df$col1)
# 原字符串无下划线时,col2会和col1内容一致,此时将col2设为NA
df$col2[df$col2 == df$col1] <- NA

df

内容的提问来源于stack exchange,提问作者Corrector

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 06:15:22