如何在R中批量处理多URL提取指定首尾词间文本
R语言批量提取学术网页指定区间文本并整合到数据框
需求说明
从dat0数据框的每个URL中,提取首个"Abstract"与末尾"Issue Section"之间的文本(不含首尾关键词),批量整合到新增的abstract列,适配700+URL的大规模处理。
实现步骤
1. 加载依赖包
需要用到网页解析包rvest、数据处理包dplyr、批量映射工具purrr和字符串处理包stringr:
library(rvest) library(dplyr) library(purrr) library(stringr)
2. 定义摘要提取函数
包含反爬延迟、错误捕获和正则匹配逻辑,确保批量处理的稳定性:
extract_abstract <- function(target_url) { # 延迟1秒,避免触发网站反爬机制 Sys.sleep(1) # 捕获网页请求错误,避免中断批量任务 page_content <- tryCatch( expr = read_html(target_url), error = function(e) return(NA) ) if (is.na(page_content)) return(NA) # 将网页转为纯文本 raw_text <- html_text(page_content) # 正则匹配目标区间文本(支持大小写不敏感) match_result <- str_match(raw_text, "(?i)abstract\\s*(.*?)\\s*(?i)issue section") # 处理匹配结果,整理格式 if (!is.na(match_result[1, 2])) { cleaned_abstract <- str_squish(match_result[1, 2]) # 恢复Background/Methods等段落的换行格式 cleaned_abstract <- str_replace_all(cleaned_abstract, "(Background|Methods|Results|Conclusions)", "\\n\\1") return(cleaned_abstract) } else { return(NA) } }
3. 批量处理生成结果
用map_chr批量应用函数,快速生成包含摘要的新数据框:
# 初始数据 dat0 <- structure(list(url = c("https://doi.org/10.1093/clinchem/hvae106.001", "https://doi.org/10.1093/clinchem/hvae106.002", "https://doi.org/10.1093/clinchem/hvae106.003" )), class = "data.frame", row.names = c(NA, -3L)) # 批量提取并整合 dat1 <- dat0 %>% mutate(abstract = map_chr(url, extract_abstract))
注意事项
- 反爬调整:如果遇到访问限制,可延长
Sys.sleep()的时间(如Sys.sleep(2)) - 正则适配:若网页中关键词格式有差异(如带冒号
Abstract:),需调整正则表达式,例如"(?i)abstract:\\s*(.*?)\\s*(?i)issue section:" - 空值处理:无法提取的URL会返回
NA,后续可通过filter(!is.na(abstract))筛选有效数据
内容的提问来源于stack exchange,提问作者denis
相关产品推荐
相关产品推荐

