在R中用rvest爬取PDF链接遇read_html连接失败,求解决方案
问题
我尝试在R语言中从网址https://providers.anthem.com/new-york-provider/claims/reimbursement-policies/爬取PDF链接,但rvest包的read_html()函数始终无响应,执行xml2::read_html(url)时出现open.connection()无法打开连接的错误。请问是否可以改用httr2来解决该问题?
原实现代码
# Load required libraries library(tidyverse) library(rvest) # Define the URL url <- "https://providers.anthem.com/new-york-provider/claims/reimbursement-policies/" # Read and process the HTML links <- try({ read_html(url) %>% html_node(xpath = "/html/body/main/div/div/div/section[3]/div/section/div[1]/section/div[1]/div/div[2]/div/p/a") %>% html_attr("href") %>% as_tibble() %>% rename(url = value) }) # Display the results with error handling if(!inherits(links, "try-error")) { print(links) } else { message("Unable to scrape the URL. This might be due to:") message("- Website requires authentication") message("- Website blocks automated scraping") message("- The XPath structure has changed") message("- Network connectivity issues") }
错误信息
> read_html(url) Error in `open.connection()`: ! cannot open the connection Hide Traceback ▆ 1. ├─xml2::read_html(url) 2. └─xml2:::read_html.default(url) 3. ├─base::suppressWarnings(...) 4. │ └─base::withCallingHandlers(...) 5. ├─xml2::read_xml(x, encoding = encoding, ..., as_html = TRUE, options = options) 6. └─xml2:::read_xml.character(...) 7. └─xml2:::read_xml.connection(...) 8. ├─base::open(x, "rb") 9. └─base::open.connection(x, "rb")
解决方案:改用httr2处理请求
可以用httr2解决这个问题。这类连接失败通常是网站拦截了默认的R请求头,httr2允许自定义请求头模拟浏览器访问,同时提供更灵活的请求控制。
修改后的代码
# 加载所需包 library(tidyverse) library(rvest) library(httr2) # 定义目标URL url <- "https://providers.anthem.com/new-york-provider/claims/reimbursement-policies/" # 使用httr2发送请求并解析HTML links <- try({ # 构建请求,添加浏览器请求头 request(url) %>% req_headers( `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", `Accept` = "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8" ) %>% req_perform() %>% # 发送请求 resp_body_html() %>% # 解析响应为HTML html_node(xpath = "/html/body/main/div/div/div/section[3]/div/section/div[1]/section/div[1]/div/div[2]/div/p/a") %>% html_attr("href") %>% as_tibble() %>% rename(url = value) }) # 结果展示与错误处理 if(!inherits(links, "try-error")) { print(links) } else { message("无法抓取URL,可能原因:") message("- 网站需要身份验证") message("- 网站拦截了自动化爬取请求") message("- 页面XPath结构已变更") message("- 网络连接问题") }
关键说明
- 自定义请求头:通过
req_headers()添加User-Agent和Accept字段,模拟普通浏览器的请求特征,避免被网站反爬机制拦截。 - 请求流程:先用httr2的
request()构建请求,req_perform()发送请求获取响应,再用resp_body_html()把响应转为rvest可处理的HTML对象,后续解析逻辑和原代码保持一致。 - 扩展处理:如果网站需要Cookie或会话验证,httr2还支持
req_cookies()、req_auth_basic()等方法处理这类场景。
内容的提问来源于stack exchange,提问作者MCP_infiltrator
相关产品推荐
相关产品推荐

