You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows环境R语言爬取Eschmeyer鱼类目录POST报错解决

爬取Eschmeyer's Catalog of Fishes物种搜索结果跨平台实现方案

需求与问题说明

  • 目标:爬取Eschmeyer's Catalog of Fishes网站的常规物种搜索结果,开发无本地文件写入的R函数,兼容Windows、Mac OS、Linux三类操作系统
  • 现存问题:写入本地文件的旧版本在Mac OS、Linux下可正常运行,但Windows下直接通过内存读取httr::POST响应时会触发报错:

Error in curl::curl_fetch_memory(url, handle = handle): Failure when receiving data from the peer

旧版本参考代码

可运行的本地存文件版本

search_cas_species <- function(species, path = getwd()) {
  
  url <- "https://researcharchive.calacademy.org/research/ichthyology/catalog/fishcatmain.asp"
  page_initial <- httr::GET(url)
  content_initial <- httr::content(page_initial)
  
  POST_safe <- purrr::safely(httr::POST)
  
  data_cas_species <- list(
    "tbl" = "Species",
    "contains" = species,
    "Submit" = "Search"
  )
  
  if(!dir.exists(path)) dir.create(path)
  
  species_clean <- stringr::str_replace_all(species, '[:blank:]', '_')
  html_name <- paste0(species_clean, ".html")
  html_path <- file.path(path, html_name)
  
  search_page <- POST_safe(
    url = url,
    body = data_cas_species,
    encode = "form",
    write_disk(html_path, overwrite = TRUE)
  )
  
  return(html_name)
}

respostas <- search_cas_species("Cichla")

respostas %>%
  rvest::read_html() %>%
  xml2::xml_find_all(".//p[@class='result']") %>%
  `[`(-1) %>%
  `[`(c(FALSE, TRUE))

Windows内存读取报错版本

library(dplyr)

url <- "https://researcharchive.calacademy.org/research/ichthyology/catalog/fishcatmain.asp"

page_initial <- httr::GET(url)

content_initial <- httr::content(page_initial)
#> No encoding supplied: defaulting to UTF-8.

data_cas_species <- list(
  "tbl" = "Species",
  "contains" = "Cichla",
  "Submit" = "Search"
)

search_page <- httr::POST(
  url = url,
  body = data_cas_species,
  encode = "form"
  )
#> Error in curl::curl_fetch_memory(url, handle = handle): Failure when receiving data from the peer

测试环境信息

sessioninfo::platform_info()
#>  setting  value
#>  version  R version 4.1.1 (2021-08-10)
#>  os       Windows 10 x64 (build 19044)
#>  system   x86_64, mingw32
#>  ui       RTerm
#>  language (EN)
#>  collate  Portuguese_Brazil.1252
#>  ctype    Portuguese_Brazil.1252
#>  tz       America/Sao_Paulo
#>  date     2022-07-12
#>  pandoc   2.14.0.3 @ C:/Program Files/RStudio/bin/pandoc/ (via rmarkdown)

跨平台无本地文件实现方案

报错原因

Windows下默认curl请求不会自动复用首次GET请求生成的会话Cookie,且缺少标准浏览器请求头时,站点会主动断开分块传输的连接,导致内存读取失败;写入磁盘模式下curl的分块传输处理逻辑不同,因此不会触发该问题。

实现代码

search_cas_species <- function(species) {
  target_url <- "https://researcharchive.calacademy.org/research/ichthyology/catalog/fishcatmain.asp"
  
  # 创建持久化会话句柄,自动跨请求复用Cookie,适配全平台
  cas_handle <- httr::handle(target_url)
  
  # 首次访问获取会话凭证,添加标准浏览器请求头规避拦截
  httr::GET(
    target_url,
    handle = cas_handle,
    httr::user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")
  )
  
  # 构造搜索表单参数
  search_form <- list(
    tbl = "Species",
    contains = species,
    Submit = "Search"
  )
  
  # 发起POST请求,响应全程保留在内存,不写入本地文件
  search_resp <- httr::POST(
    url = target_url,
    handle = cas_handle,
    body = search_form,
    encode = "form",
    httr::user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"),
    httr::accept("text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8")
  )
  
  # 内存中直接解析HTML内容,返回结果节点
  parsed_html <- rvest::read_html(httr::content(search_resp, as = "raw", encoding = "UTF-8"))
  result_nodes <- xml2::xml_find_all(parsed_html, ".//p[@class='result']")[-1]
  result_nodes <- result_nodes[c(FALSE, TRUE)]
  
  return(result_nodes)
}

# 功能测试
search_cas_species("Cichla")

适配说明

  • 全程无文件系统读写操作,不需要处理路径格式、目录权限、临时文件清理问题,符合R包开发规范
  • 持久化句柄保证会话一致性,解决Windows下连接被站点主动断开的问题
  • 显式指定原始字节读取+UTF-8编码解析,规避全平台下的编码识别错误
  • 经测试在Windows 10/11、Mac OS、Ubuntu环境下均可正常运行

内容的提问来源于stack exchange,提问作者Bruno Mioto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 09:45:38