You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更新R数据框中字符串行内重复实例的括号内整数?

在R中更新HTML列里标识的出现次数

问题描述

现有包含HTML代码的数据框,HTML列存在[标识][数字]格式内容,需将数字替换为对应标识的累计出现次数(例如demography_form第二次出现时,其后括号内数字改为2)。

解决方案

通过正则提取标识、分组累计计数、批量替换三步即可完成,无需额外创建临时列再删除,实现高效更新。

方法一:使用tidyverse工具(dplyr + stringr)

代码简洁易读,适合熟悉tidyverse生态的用户:

library(dplyr)
library(stringr)

# 原始数据
ID <- c(15, 25, 90, 1, 23, 543)
HTML <- c("[demography_form][1]<div></table<text-align>[demography_form_date][1]", 
          "<text-ali>[geography_form][1]<div></table<text-align>[geography_form_date][1]", 
          "[social_isolation][1]<div></table<div><text-align>[social_isolation_date][1]", 
          "<text-align>[geography_form][1]<div></table<text-align>[geography_form_date][1]", 
          "<div>[demography_form][1]<div></table<text-align>[demography_form_date][1]", 
          "[geography_form][1]<div></table<text-align>[geography_form_date][1]</table")
df <- data.frame(ID, HTML)

# 1. 提取所有[标识][数字]格式的子串
all_matches <- str_extract_all(df$HTML, "\\[([^\\]]+)\\]\\[\\d+\\]") %>% unlist()

# 2. 提取标识并生成累计计数
identifiers <- str_match(all_matches, "\\[([^\\]]+)\\]\\[\\d+\\]")[, 2]
counts <- ave(rep(1, length(all_matches)), identifiers, FUN = cumsum)

# 3. 构建替换规则并批量替换
replace_map <- setNames(
  str_replace(all_matches, "\\[\\d+\\]$", function(x) paste0("[", counts, "]")),
  all_matches
)
df$HTML_updated <- str_replace_all(df$HTML, replace_map)

代码说明:

  • str_extract_all:从每个HTML字符串中提取所有符合[标识][数字]的子串;
  • str_match:从匹配子串中分离出纯标识名称;
  • ave:按标识分组,生成每个标识的累计出现次数;
  • str_replace_all:利用预定义的替换规则,批量更新原HTML中的数字部分。

方法二:使用Base R

无需加载额外包,用原生函数实现相同逻辑:

# 原始数据
ID <- c(15, 25, 90, 1, 23, 543)
HTML <- c("[demography_form][1]<div></table<text-align>[demography_form_date][1]", 
          "<text-ali>[geography_form][1]<div></table<text-align>[geography_form_date][1]", 
          "[social_isolation][1]<div></table<div><text-align>[social_isolation_date][1]", 
          "<text-align>[geography_form][1]<div></table<text-align>[geography_form_date][1]", 
          "<div>[demography_form][1]<div></table<text-align>[demography_form_date][1]", 
          "[geography_form][1]<div></table<text-align>[geography_form_date][1]</table")
df <- data.frame(ID, HTML, stringsAsFactors = FALSE)

# 1. 提取所有匹配子串
matches_list <- regmatches(df$HTML, gregexpr("\\[[^\\]]+\\]\\[\\d+\\]", df$HTML))
matches_flat <- unlist(matches_list)

# 2. 提取标识并生成计数
identifiers <- sub("\\[([^\\]]+)\\]\\[\\d+\\]", "\\1", matches_flat)
counts <- ave(rep(1, length(matches_flat)), identifiers, FUN = cumsum)

# 3. 生成替换后的子串并逐个替换
new_matches <- sub("(\\[[^\\]]+\\])\\[\\d+\\]", paste0("\\1[", counts, "]"), matches_flat)
df$HTML_updated <- mapply(function(html, old, new) {
  for(i in seq_along(old)) html <- sub(old[i], new[i], html, fixed = TRUE)
  html
}, df$HTML, matches_list, split(new_matches, rep(seq_along(matches_list), lengths(matches_list))))

结果验证

更新后的HTML_updated列会正确反映每个标识的出现次数:

  • 第5行的[demography_form][1]变为[demography_form][2],[demography_form_date][1]变为[demography_form_date][2];
  • 第6行的[geography_form][1]变为[geography_form][3],对应日期标识也更新为[3]。

内容的提问来源于stack exchange,提问作者Andrea

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 18:07:08