You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:保留指定字符及后续数字并在数据框生成新列

解决方案:提取指定前缀及后续数字并生成新列

需求说明

处理字符串向量,仅保留指定字符集合C_Keep中的前缀,以及紧跟在这些前缀后的数字,最终将处理结果作为新列添加到数据框中。


使用stringr包实现(简洁高效)

stringr是R中常用的字符串处理工具包,语法直观易读。

library(stringr)

# 定义需要保留的前缀集合
C_Keep <- c("ab", "acr", "sb", "scr")

# 构建正则匹配模式:匹配完整前缀 + 后续的空格+数字组合
pattern <- paste0("\\b(", paste(C_Keep, collapse = "|"), ")\\b(\\s+\\d+)*")

# 示例数据框
text <- "ab 187; acr 76 98 298 876 987; legislature governors office attorney generals office re gaming issues ab 1416 165 187 267; calepa support sb 265; scr 17689 986 83783 3982"
df <- data.frame(text)

# 定义处理函数
extract_targets <- function(x) {
  # 按分号分割字符串为多个片段
  segments <- str_split(x, ";\\s*")[[1]]
  # 提取每个片段中符合模式的内容并拼接
  processed_segs <- sapply(segments, function(seg) {
    matches <- str_extract_all(seg, pattern)[[1]]
    paste(matches, collapse = " ")
  })
  # 过滤空片段后重新拼接成结果字符串
  paste(processed_segs[processed_segs != ""], collapse = "; ")
}

# 添加output列到数据框
df$output <- sapply(df$text, extract_targets)

# 查看结果
print(df)

代码解释

  • 正则模式:\\b确保匹配完整的前缀(避免匹配其他单词中的子串,比如"abc"里的"ab");(\\s+\\d+)*匹配前缀后任意数量的"空格+数字"组合。
  • 处理逻辑:先将原字符串按分号分割为独立片段,对每个片段提取目标内容,最后过滤空片段并重新拼接成符合要求的字符串。

使用Base R实现(无需额外包)

如果不想安装第三方包,可以用Base R的字符串函数实现相同功能:

# 定义需要保留的前缀集合
C_Keep <- c("ab", "acr", "sb", "scr")
pattern <- paste0("\\b(", paste(C_Keep, collapse = "|"), ")\\b(\\s+\\d+)*")

# 示例数据框
text <- "ab 187; acr 76 98 298 876 987; legislature governors office attorney generals office re gaming issues ab 1416 165 187 267; calepa support sb 265; scr 17689 986 83783 3982"
df <- data.frame(text)

# 定义处理函数
extract_targets_base <- function(x) {
  segments <- strsplit(x, ";\\s*")[[1]]
  processed_segs <- sapply(segments, function(seg) {
    matches <- regmatches(seg, gregexpr(pattern, seg))[[1]]
    paste(matches, collapse = " ")
  })
  paste(processed_segs[processed_segs != ""], collapse = "; ")
}

# 添加output列
df$output <- sapply(df$text, extract_targets_base)

# 查看结果
print(df)

代码解释

用Base R的strsplit替代str_split,用gregexpr+regmatches组合替代str_extract_all,核心逻辑与stringr版本完全一致。


输出结果

运行代码后,数据框df的output列将与需求中的期望结果完全一致:

text
1 ab 187; acr 76 98 298 876 987; legislature governors office attorney generals office re gaming issues ab 1416 165 187 267; calepa support sb 265; scr 17689 986 83783 3982
                                                                                              output
1                        ab 187; acr 76 98 298 876 987; ab 1416 165 187 267; sb 265; scr 17689 986 83783 3982

内容的提问来源于stack exchange,提问作者Sharif Amlani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 15:25:19