如何更新R数据框中字符串行内重复实例的括号内整数?
在R中更新HTML列里标识的出现次数
问题描述
现有包含HTML代码的数据框,HTML列存在[标识][数字]格式内容,需将数字替换为对应标识的累计出现次数(例如demography_form第二次出现时,其后括号内数字改为2)。
解决方案
通过正则提取标识、分组累计计数、批量替换三步即可完成,无需额外创建临时列再删除,实现高效更新。
方法一:使用tidyverse工具(dplyr + stringr)
代码简洁易读,适合熟悉tidyverse生态的用户:
library(dplyr) library(stringr) # 原始数据 ID <- c(15, 25, 90, 1, 23, 543) HTML <- c("[demography_form][1]<div></table<text-align>[demography_form_date][1]", "<text-ali>[geography_form][1]<div></table<text-align>[geography_form_date][1]", "[social_isolation][1]<div></table<div><text-align>[social_isolation_date][1]", "<text-align>[geography_form][1]<div></table<text-align>[geography_form_date][1]", "<div>[demography_form][1]<div></table<text-align>[demography_form_date][1]", "[geography_form][1]<div></table<text-align>[geography_form_date][1]</table") df <- data.frame(ID, HTML) # 1. 提取所有[标识][数字]格式的子串 all_matches <- str_extract_all(df$HTML, "\\[([^\\]]+)\\]\\[\\d+\\]") %>% unlist() # 2. 提取标识并生成累计计数 identifiers <- str_match(all_matches, "\\[([^\\]]+)\\]\\[\\d+\\]")[, 2] counts <- ave(rep(1, length(all_matches)), identifiers, FUN = cumsum) # 3. 构建替换规则并批量替换 replace_map <- setNames( str_replace(all_matches, "\\[\\d+\\]$", function(x) paste0("[", counts, "]")), all_matches ) df$HTML_updated <- str_replace_all(df$HTML, replace_map)
代码说明:
str_extract_all:从每个HTML字符串中提取所有符合[标识][数字]的子串;str_match:从匹配子串中分离出纯标识名称;ave:按标识分组,生成每个标识的累计出现次数;str_replace_all:利用预定义的替换规则,批量更新原HTML中的数字部分。
方法二:使用Base R
无需加载额外包,用原生函数实现相同逻辑:
# 原始数据 ID <- c(15, 25, 90, 1, 23, 543) HTML <- c("[demography_form][1]<div></table<text-align>[demography_form_date][1]", "<text-ali>[geography_form][1]<div></table<text-align>[geography_form_date][1]", "[social_isolation][1]<div></table<div><text-align>[social_isolation_date][1]", "<text-align>[geography_form][1]<div></table<text-align>[geography_form_date][1]", "<div>[demography_form][1]<div></table<text-align>[demography_form_date][1]", "[geography_form][1]<div></table<text-align>[geography_form_date][1]</table") df <- data.frame(ID, HTML, stringsAsFactors = FALSE) # 1. 提取所有匹配子串 matches_list <- regmatches(df$HTML, gregexpr("\\[[^\\]]+\\]\\[\\d+\\]", df$HTML)) matches_flat <- unlist(matches_list) # 2. 提取标识并生成计数 identifiers <- sub("\\[([^\\]]+)\\]\\[\\d+\\]", "\\1", matches_flat) counts <- ave(rep(1, length(matches_flat)), identifiers, FUN = cumsum) # 3. 生成替换后的子串并逐个替换 new_matches <- sub("(\\[[^\\]]+\\])\\[\\d+\\]", paste0("\\1[", counts, "]"), matches_flat) df$HTML_updated <- mapply(function(html, old, new) { for(i in seq_along(old)) html <- sub(old[i], new[i], html, fixed = TRUE) html }, df$HTML, matches_list, split(new_matches, rep(seq_along(matches_list), lengths(matches_list))))
结果验证
更新后的HTML_updated列会正确反映每个标识的出现次数:
- 第5行的
[demography_form][1]变为[demography_form][2],[demography_form_date][1]变为[demography_form_date][2]; - 第6行的
[geography_form][1]变为[geography_form][3],对应日期标识也更新为[3]。
内容的提问来源于stack exchange,提问作者Andrea
相关产品推荐
相关产品推荐

