You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中用rvest爬取PDF链接遇read_html连接失败,求解决方案

问题

我尝试在R语言中从网址https://providers.anthem.com/new-york-provider/claims/reimbursement-policies/爬取PDF链接,但rvest包的read_html()函数始终无响应,执行xml2::read_html(url)时出现open.connection()无法打开连接的错误。请问是否可以改用httr2来解决该问题?

原实现代码

# Load required libraries
library(tidyverse)
library(rvest)

# Define the URL
url <- "https://providers.anthem.com/new-york-provider/claims/reimbursement-policies/"

# Read and process the HTML
links <- try({
  read_html(url) %>%
    html_node(xpath = "/html/body/main/div/div/div/section[3]/div/section/div[1]/section/div[1]/div/div[2]/div/p/a") %>%
    html_attr("href") %>%
    as_tibble() %>%
    rename(url = value)
})

# Display the results with error handling
if(!inherits(links, "try-error")) {
  print(links)
} else {
  message("Unable to scrape the URL. This might be due to:")
  message("- Website requires authentication")
  message("- Website blocks automated scraping")
  message("- The XPath structure has changed")
  message("- Network connectivity issues")
}

错误信息

>   read_html(url)
Error in `open.connection()`:
! cannot open the connection
Hide Traceback
    ▆
 1. ├─xml2::read_html(url)
 2. └─xml2:::read_html.default(url)
 3.   ├─base::suppressWarnings(...)
 4.   │ └─base::withCallingHandlers(...)
 5.   ├─xml2::read_xml(x, encoding = encoding, ..., as_html = TRUE, options = options)
 6.   └─xml2:::read_xml.character(...)
 7.     └─xml2:::read_xml.connection(...)
 8.       ├─base::open(x, "rb")
 9.       └─base::open.connection(x, "rb")

解决方案:改用httr2处理请求

可以用httr2解决这个问题。这类连接失败通常是网站拦截了默认的R请求头,httr2允许自定义请求头模拟浏览器访问,同时提供更灵活的请求控制。

修改后的代码

# 加载所需包
library(tidyverse)
library(rvest)
library(httr2)

# 定义目标URL
url <- "https://providers.anthem.com/new-york-provider/claims/reimbursement-policies/"

# 使用httr2发送请求并解析HTML
links <- try({
  # 构建请求,添加浏览器请求头
  request(url) %>%
    req_headers(
      `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
      `Accept` = "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8"
    ) %>%
    req_perform() %>%  # 发送请求
    resp_body_html() %>%  # 解析响应为HTML
    html_node(xpath = "/html/body/main/div/div/div/section[3]/div/section/div[1]/section/div[1]/div/div[2]/div/p/a") %>%
    html_attr("href") %>%
    as_tibble() %>%
    rename(url = value)
})

# 结果展示与错误处理
if(!inherits(links, "try-error")) {
  print(links)
} else {
  message("无法抓取URL,可能原因:")
  message("- 网站需要身份验证")
  message("- 网站拦截了自动化爬取请求")
  message("- 页面XPath结构已变更")
  message("- 网络连接问题")
}

关键说明

  • 自定义请求头:通过req_headers()添加User-Agent和Accept字段,模拟普通浏览器的请求特征,避免被网站反爬机制拦截。
  • 请求流程:先用httr2的request()构建请求,req_perform()发送请求获取响应,再用resp_body_html()把响应转为rvest可处理的HTML对象,后续解析逻辑和原代码保持一致。
  • 扩展处理:如果网站需要Cookie或会话验证,httr2还支持req_cookies()、req_auth_basic()等方法处理这类场景。

内容的提问来源于stack exchange,提问作者MCP_infiltrator

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 10:03:24