You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言DataFrame中将指定子串拼接为新列?

提取DataFrame列子串并拼接新列的实现方法

你需要从Combination列提取第一个下划线之前的子串,再与new_var列内容拼接生成新列,以下是两种在R中实现的方案:

原始数据

df = structure(list(Combination = c("Animalborne_Archival", "Animalborne_Satellite_Archival", 
"BRUV_Acoustic", "Stationary_Acoustic_Radio_PIT", "Stationary_Acoustic", 
"Controlled_Archival"), new_var = c(Animalborne_Archival = "NoReceiver", 
Animalborne_Satellite_Archival = "NoReceiver", BRUV_Acoustic = "Receiver", 
Stationary_Acoustic_Radio_PIT = "Receiver", Stationary_Acoustic = "Receiver", 
Controlled_Archival = "NoReceiver")), row.names = c(7L, 188L, 
154L, 41L, 134L, 159L), class = "data.frame")

方案1:Base R 原生实现

使用sub()函数匹配并移除第一个下划线及之后的所有内容,再用paste()完成拼接:

# 提取Combination列第一个下划线前的子串,拼接后生成新列new_col
df$new_col <- paste(sub("_.*", "", df$Combination), df$new_var, sep = "_")
  • sub("_.*", "", df$Combination):正则表达式_.*匹配第一个下划线及后续所有字符,替换为空字符串,得到目标子串
  • sep = "_":指定拼接的分隔符,可根据需求改为空格或其他符号

方案2:Tidyverse 工具链实现

借助dplyr的列操作和stringr的字符串处理函数:

library(dplyr)
library(stringr)

df <- df %>%
  mutate(new_col = str_c(
    str_extract(Combination, "^[^_]+"),  # 提取开头到第一个下划线前的内容
    new_var,
    sep = "_"
  ))
  • str_extract(Combination, "^[^_]+"):正则^[^_]+匹配从字符串开头到第一个下划线的所有非下划线字符
  • str_c()相当于tidyverse版本的paste(),用法更简洁

处理后结果示例

生成的new_col列内容如下:

[1] "Animalborne_NoReceiver" "Animalborne_NoReceiver" "BRUV_Receiver"         
[4] "Stationary_Receiver"    "Stationary_Receiver"    "Controlled_NoReceiver" 

内容的提问来源于stack exchange,提问作者Kristen Cyr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 09:07:32