You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中放慢read_html请求速度以规避HTTP 429错误?

网页爬取反爬规避方案

问题背景

使用rvest爬取某论坛207页帖子时,触发反爬机制报错:Error in open.connection(x, "rb") : HTTP error 429。手动单页请求可临时恢复,但网站限制10秒内多页访问,需实现每5秒请求一页的访问间隔。现有代码如下:

url_base <- "https:// website here"

map_df(1:207, function(i) {
  
  # simple but effective progress indicator
  cat(".")
  
  pg <- read_html(sprintf(url_base, i))
  
  data.frame(text=html_text(html_nodes(pg, ".text")),
             date=html_text(html_nodes(pg, "time")),
             coffee=html_text(html_nodes(pg, ".coffeetyp_3sch")),
             stringsAsFactors=FALSE)
  
}) -> posts

实现方法

基础延迟方案

在每次请求后添加Sys.sleep(5)实现固定访问间隔,修改后的代码如下:

url_base <- "https:// website here"

map_df(1:207, function(i) {
  
  cat(".")
  
  pg <- read_html(sprintf(url_base, i))
  
  # 请求完成后休眠5秒,避免短时间内频繁访问
  Sys.sleep(5)
  
  data.frame(text=html_text(html_nodes(pg, ".text")),
             date=html_text(html_nodes(pg, "time")),
             coffee=html_text(html_nodes(pg, ".coffeetyp_3sch")),
             stringsAsFactors=FALSE)
  
}) -> posts

优化版延迟(跳过首次休眠)

如果不想在第一页请求前浪费时间,可仅在非首次请求前添加休眠:

url_base <- "https:// website here"

map_df(1:207, function(i) {
  
  # 除第一页外,每次请求前休眠5秒
  if (i != 1) {
    Sys.sleep(5)
  }
  
  cat(".")
  
  pg <- read_html(sprintf(url_base, i))
  
  data.frame(text=html_text(html_nodes(pg, ".text")),
             date=html_text(html_nodes(pg, "time")),
             coffee=html_text(html_nodes(pg, ".coffeetyp_3sch")),
             stringsAsFactors=FALSE)
  
}) -> posts

进阶反爬优化

为进一步降低被识别风险,可做以下调整:

  • 随机化休眠时间:用Sys.sleep(runif(1, 3, 7))替代固定5秒,模拟人类访问的随机性
  • 异常重试机制:用tryCatch捕获HTTP 429错误,遇到时延长休眠时间后重试,避免程序中断

内容的提问来源于stack exchange,提问作者Buckethead

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 14:05:38