You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中使用gsub正则提取列内多段引号整数值并求和的实现方法

R数据框字符串列提取数值并求和实现方案

现有代码问题

你当前使用的gsub()正则规则仅能匹配到第一个被双引号包裹的数值,剩余字符未被正确匹配替换,既无法插入分隔符也会残留多余符号。

实现方案

方案1:stringr包简便实现(推荐)

直接提取所有符合规则的数值,无需手动处理分隔符,一步完成求和:

library(stringr)
# 直接生成求和结果列
proto_runs$sum_value <- sapply(
  str_extract_all(proto_runs$CompletedCycles, '(?<=")\\d+(?=")'),
  function(x) sum(as.integer(x))
)

如果需要先得到空格分隔的数值字符串,可额外增加一行代码:

proto_runs$value_str <- sapply(
  str_extract_all(proto_runs$CompletedCycles, '(?<=")\\d+(?=")'),
  paste, collapse = " "
)

正则说明:(?<=")\\d+(?=")使用前后断言匹配被双引号包裹的纯数字,自动过滤所有非目标内容。

方案2:纯Base R实现(无需额外安装包)

使用regmatches + gregexpr完成匹配提取:

# 匹配所有符合规则的数值
match_res <- gregexpr('(?<=")\\d+(?=")', proto_runs$CompletedCycles, perl = TRUE)
# 提取数值后求和
proto_runs$sum_value <- sapply(
  regmatches(proto_runs$CompletedCycles, match_res),
  function(x) sum(as.integer(x))
)

如果需要空格分隔的字符串:

proto_runs$value_str <- sapply(
  regmatches(proto_runs$CompletedCycles, match_res),
  paste, collapse = " "
)

方案3:仅用gsub实现

如果要求必须使用gsub完成,可通过两次替换生成目标格式:

# 替换所有非数字、非双引号的字符为空,再把引号替换为空格,最后去掉首尾空格
proto_runs$value_str <- trimws(gsub('"', ' ', gsub('[^\\d"]', '', proto_runs$CompletedCycles)))
# 求和
proto_runs$sum_value <- sapply(
  strsplit(proto_runs$value_str, " "),
  function(x) sum(as.integer(x))
)

内容的提问来源于stack exchange,提问作者Christopher Morrissey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 03:54:01