You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R代码无法从子页面提取通讯作者元素的问题求助

改进R代码以提取通讯作者信息

问题分析

你当前的代码能成功获取文章子页面,但无法定位到通讯作者的<a id="corresp-c1">标签,核心原因包括:

  • 部分页面的通讯作者标签ID并非固定为corresp-c1(多通讯作者场景下ID会动态变化)
  • 手动处理httr响应的解析方式不够稳健,易引入编码或格式问题
  • 选择器过于刚性,未考虑标签的上下文嵌套结构

代码改进建议(保留原辅助函数)

以下是基于原代码结构的优化方案,核心是优化选择器、页面读取逻辑,并添加必要的容错处理:

  1. 优化页面读取函数
    直接使用rvest::read_html()处理URL,无需手动调用httr::GET和content,该函数会自动处理HTTP请求和HTML解析,更简洁可靠。

  2. 采用灵活的CSS选择器
    放弃固定ID的选择器,改用:

    • 先定位通讯作者的父容器(比如.correspondence-author类),再查找内部的邮件链接
    • 或使用属性选择器a[href^="mailto:"]结合上下文筛选通讯作者(避免误抓其他邮件链接)
  3. 添加错误处理与信息提取
    在提取函数中加入容错机制,避免单个页面失败导致整个流程中断,同时从找到的标签中分别提取姓名和邮箱信息。

完整改进代码

library(rvest) 
library(tidyverse)

# 保留原辅助函数:拼接子页面URL
merge_strings <- function(x){
  prefix_str_1 <- "https://molecularbrain.biomedcentral.com/"
  paste0(prefix_str_1, x)
}

# 优化页面读取函数:直接用read_html处理URL,添加容错
read_page_1 <- function(x){
  purrr::possibly(rvest::read_html, otherwise = NULL)(x)
}

# 改进通讯作者提取函数
correspondence_search <- function(html_node){
  if(is.null(html_node)) return(list(name = NA, email = NA))
  
  # 优先通过通讯作者容器定位,精准抓取目标链接
  corresp_node <- html_node(html_node, ".correspondence-author a[href^='mailto:']")
  
  # 兼容不同ID格式的情况
  if(is.na(corresp_node)){
    corresp_node <- html_node(html_node, "a[id^='corresp-']")
  }
  
  # 提取姓名和邮箱并清理格式
  name <- html_text(corresp_node, trim = TRUE)
  email <- str_remove(html_attr(corresp_node, "href"), "^mailto:")
  
  return(list(name = name, email = email))
}

# 主流程:获取文章列表与子页面链接
str_1 <- "https://molecularbrain.biomedcentral.com/articles"
html <- read_html(str_1)

c_listing_title <- html_elements(html,"h3.c-listing__title")
a_element <- html_node(c_listing_title,"a")
a_href <- as.list(html_attr(a_element,"href"))
sub_pages <- lapply(a_href, merge_strings)

# 加载所有子页面
collection_html_sub_pages <- lapply(sub_pages, read_page_1)

# 提取所有通讯作者信息
correspondence_authors <- lapply(collection_html_sub_pages, correspondence_search)

# 转换为数据框方便后续处理
corresp_df <- bind_rows(correspondence_authors)
print(corresp_df)

关键说明

  • 选择器优先级:先通过.correspondence-author容器定位,确保只抓取通讯作者的联系方式,避免误抓其他作者的邮件链接
  • 容错机制:purrr::possibly保证单个页面加载失败时,流程不会中断,而是返回NA值
  • 信息拆分:将姓名和邮箱分开提取,便于后续的数据分析或存储

内容的提问来源于stack exchange,提问作者Jose Serra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 07:15:08