使用rvest爬取网站时遇“cannot open the connection”错误求排查
无法建立连接:
Error in open.connection(x, "rb") : cannot open the connection 我用R的tidyverse和rvest编写了爬虫代码,尝试获取https://tmarketonline.bg的蔬果类商品信息,但运行时出现上述连接错误:
library(tidyverse) library(rvest) fruits <- read_html("https://tmarketonline.bg/category/plodove-zelenchuci-i-yadki?page=1") fruits_df <- fruits %>% html_elements("._product") %>% map_dfr(~ tibble( product = .x %>% html_element("._product-name-tag a") %>% html_text2(), price = .x %>% html_element("._product-price-inner span") %>% html_text2(), price_old = .x %>% html_element("._product-price-old") %>% html_text2(), unit = .x %>% html_element("._button_unit") %>% html_text2())) %>% mutate(date = Sys.Date(), type = "Fruits and vegetables", source = "T MARKET", .before = product) %>% mutate(price = str_replace(price, ",", "."), price = parse_number(price)) %>% distinct()
以下是排查线索:
验证网站可达性
直接用浏览器打开目标URL,如果无法加载,说明是网站本身不可用、你的网络有访问限制(比如地区封锁);如果浏览器能正常打开,再继续排查。添加请求头绕过反爬
很多网站会拦截无浏览器标识的请求,试试给请求添加标准User-Agent头:library(httr) # 模拟Chrome浏览器请求头 headers <- c( "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" ) # 先发送带请求头的请求 response <- GET("https://tmarketonline.bg/category/plodove-zelenchuci-i-yadki?page=1", add_headers(.headers = headers)) # 再解析响应 fruits <- read_html(response)如果还是失败,可以从浏览器开发者工具中复制完整请求头(含Cookie、Referer等)补充到请求里。
检查网络依赖和包版本
- 运行
curl::has_internet(),若返回FALSE,说明R无法连接互联网,检查系统网络设置。 - 更新相关包到最新版本:
update.packages(c("rvest", "httr", "tidyverse")),旧版本可能存在兼容性问题。
- 运行
配置代理(若网络需代理)
如果你所在网络需要代理才能访问目标网站,设置系统代理环境变量:Sys.setenv(HTTP_PROXY = "http://你的代理地址:端口") Sys.setenv(HTTPS_PROXY = "http://你的代理地址:端口")
内容的提问来源于stack exchange,提问作者NickD
相关产品推荐
相关产品推荐

