You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R从PubMed子页面抓取DOI、被引次数等多类数据?

扩展PubMed数据抓取R代码:获取DOI、被引次数等更多信息

原代码已实现从PubMed搜索页抓取文献基础信息,并通过子页面获取摘要。以下是扩展版本,可同时抓取DOI、被引次数、期刊名称、发表时间等更多字段:

修改后的完整代码

library(tidyverse)
library(rvest)

# 读取PubMed搜索页面
page <- "https://pubmed.ncbi.nlm.nih.gov/?term=((((((%E2%80%98Food%20Supply%E2%80%99%20(MeSH))%20OR%20%E2%80%98Food%20Storage%E2%80%99%20(MeSH))%20OR%20%E2%80%98Hunger%E2%80%99(MeSH)%20OR%20food%20security%20OR%20food%20insecurity%20OR%20household%20food%20security%20OR%20global%20food%20security)%20OR%20household%20food%20insecurity)))%20AND%20((%E2%80%98Prevalence%E2%80%99%20(MeSH))%20OR%20%E2%80%98Cross-Sectional%20Studies%E2%80%99%20(MeSH)%20OR%20cross-sectional%20study%20OR%20Prevalence%20Studies%20OR%20prevalence%20study%20OR%20Cross-Sectional%20Analyses%20OR%20CrossSectional%20Analysis%20OR%20Cross%20Sectional%20Analysis%20OR%20Cross%20Sectional%20Analyses)&filter=lang.english&filter=lang.portuguese" %>% 
  read_html()

# 提取搜索页的基础文献信息
df <- page %>% 
  html_elements(".docsum-content") %>% 
  map_dfr(~ tibble(
    title = .x %>% html_element(".docsum-title") %>% html_text2(),
    authors = .x %>% html_element(".full-authors") %>% html_text2(),
    PMID = .x %>% html_element(".docsum-pmid") %>% html_text2(),
    synopsis = .x %>% html_element(".full-view-snippet") %>% html_text2(),
    link = .x %>% html_element(".docsum-title") %>% html_attr("href") %>% str_c("https://pubmed.ncbi.nlm.nih.gov", .)
  ))

# 定义函数:从详情页抓取多个字段
get_pubmed_details <- function(link) {
  cat("正在抓取:", link, "\n")
  # 添加延迟,避免请求过于频繁被拦截
  Sys.sleep(1)
  
  page_details <- read_html(link)
  
  tibble(
    abstract = page_details %>% html_element(".abstract-content.selected") %>% html_text2(),
    doi = page_details %>% html_element(".identifiers .doi a") %>% html_text2(),
    citation_count = page_details %>% html_element(".article-citations .count") %>% html_text2() %>% as.integer(),
    journal = page_details %>% html_element(".journal-actions-trigger") %>% html_text2(),
    pub_date = page_details %>% html_element(".cit") %>% html_text2() %>% str_extract("\\d{4}.*$")
  )
}

# 整合所有数据
df_full <- df %>% 
  bind_cols(map_dfr(.$link, get_pubmed_details))

# 查看结果
head(df_full)

关键说明

  1. 统一字段抓取函数:将原有的get_abstract替换为get_pubmed_details,一次请求即可获取多个字段,减少HTTP请求次数,降低被网站拦截的风险。
  2. 字段选择器说明:
    • DOI:通过.identifiers .doi a定位详情页的DOI链接文本
    • 被引次数:通过.article-citations .count获取被引数字,转换为整数类型
    • 期刊名称:通过.journal-actions-trigger提取期刊名
    • 发表时间:从.cit元素中提取日期部分
  3. 防拦截处理:添加Sys.sleep(1)在每次请求后延迟1秒,避免触发网站的反爬机制。
  4. 空值处理:若某字段不存在(如无被引次数),html_text2()会返回NA,不影响整体数据整合。

内容的提问来源于stack exchange,提问作者Helena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 21:11:29