如何用R从PubMed子页面抓取DOI、被引次数等多类数据?
扩展PubMed数据抓取R代码:获取DOI、被引次数等更多信息
原代码已实现从PubMed搜索页抓取文献基础信息,并通过子页面获取摘要。以下是扩展版本,可同时抓取DOI、被引次数、期刊名称、发表时间等更多字段:
修改后的完整代码
library(tidyverse) library(rvest) # 读取PubMed搜索页面 page <- "https://pubmed.ncbi.nlm.nih.gov/?term=((((((%E2%80%98Food%20Supply%E2%80%99%20(MeSH))%20OR%20%E2%80%98Food%20Storage%E2%80%99%20(MeSH))%20OR%20%E2%80%98Hunger%E2%80%99(MeSH)%20OR%20food%20security%20OR%20food%20insecurity%20OR%20household%20food%20security%20OR%20global%20food%20security)%20OR%20household%20food%20insecurity)))%20AND%20((%E2%80%98Prevalence%E2%80%99%20(MeSH))%20OR%20%E2%80%98Cross-Sectional%20Studies%E2%80%99%20(MeSH)%20OR%20cross-sectional%20study%20OR%20Prevalence%20Studies%20OR%20prevalence%20study%20OR%20Cross-Sectional%20Analyses%20OR%20CrossSectional%20Analysis%20OR%20Cross%20Sectional%20Analysis%20OR%20Cross%20Sectional%20Analyses)&filter=lang.english&filter=lang.portuguese" %>% read_html() # 提取搜索页的基础文献信息 df <- page %>% html_elements(".docsum-content") %>% map_dfr(~ tibble( title = .x %>% html_element(".docsum-title") %>% html_text2(), authors = .x %>% html_element(".full-authors") %>% html_text2(), PMID = .x %>% html_element(".docsum-pmid") %>% html_text2(), synopsis = .x %>% html_element(".full-view-snippet") %>% html_text2(), link = .x %>% html_element(".docsum-title") %>% html_attr("href") %>% str_c("https://pubmed.ncbi.nlm.nih.gov", .) )) # 定义函数:从详情页抓取多个字段 get_pubmed_details <- function(link) { cat("正在抓取:", link, "\n") # 添加延迟,避免请求过于频繁被拦截 Sys.sleep(1) page_details <- read_html(link) tibble( abstract = page_details %>% html_element(".abstract-content.selected") %>% html_text2(), doi = page_details %>% html_element(".identifiers .doi a") %>% html_text2(), citation_count = page_details %>% html_element(".article-citations .count") %>% html_text2() %>% as.integer(), journal = page_details %>% html_element(".journal-actions-trigger") %>% html_text2(), pub_date = page_details %>% html_element(".cit") %>% html_text2() %>% str_extract("\\d{4}.*$") ) } # 整合所有数据 df_full <- df %>% bind_cols(map_dfr(.$link, get_pubmed_details)) # 查看结果 head(df_full)
关键说明
- 统一字段抓取函数:将原有的
get_abstract替换为get_pubmed_details,一次请求即可获取多个字段,减少HTTP请求次数,降低被网站拦截的风险。 - 字段选择器说明:
- DOI:通过
.identifiers .doi a定位详情页的DOI链接文本 - 被引次数:通过
.article-citations .count获取被引数字,转换为整数类型 - 期刊名称:通过
.journal-actions-trigger提取期刊名 - 发表时间:从
.cit元素中提取日期部分
- DOI:通过
- 防拦截处理:添加
Sys.sleep(1)在每次请求后延迟1秒,避免触发网站的反爬机制。 - 空值处理:若某字段不存在(如无被引次数),
html_text2()会返回NA,不影响整体数据整合。
内容的提问来源于stack exchange,提问作者Helena
相关产品推荐
相关产品推荐

