使用Sys.sleep()仍无法解决爬虫HTTP 429请求过多错误
解决爬虫HTTP 429错误及函数拆分问题
核心问题分析
当前代码的最大问题是每个车辆属性函数都重复请求同一车辆页面,比如提取颜色、变速箱、马力等属性时,每个函数都会单独调用read_html(car_link),导致单辆车的页面被请求近10次,这直接触发了服务器的请求频率限制,即使加了全局延迟也无济于事。
解决方案
1. 每个车辆页面仅请求一次,拆分属性提取逻辑
保留属性拆分的函数,但修改为接收已获取的页面内容,而非重新请求链接:
library("robotstxt") library("dplyr") library("rvest") library("purrr") # 初始化空数据框 CARS <- data.frame() # -------------------------- # 定义页面获取与属性提取函数 # -------------------------- # 仅请求一次车辆页面 get_car_page <- function(car_link) { read_html(car_link) } # 提取颜色(参数为已获取的页面) get_color <- function(car_page) { car_page %>% html_node("tr:nth-child(12) span") %>% html_text2() } # 提取变速箱系统 get_gearing_sys <- function(car_page) { car_page %>% html_node("tr:nth-child(11) span") %>% html_text2() } # 提取马力 get_car_HP <- function(car_page) { car_page %>% html_node("tr:nth-child(10) span") %>% html_text2() } # 提取排量 get_car_CC <- function(car_page) { car_page %>% html_node("tr:nth-child(9) span") %>% html_text2() } # 提取燃油类型 get_car_fuel <- function(car_page) { car_page %>% html_node("tr:nth-child(8) span") %>% html_text2() } # 提取里程 get_car_km <- function(car_page) { car_page %>% html_node("tr:nth-child(7) span") %>% html_text2() } # 提取上牌日期 get_car_date <- function(car_page) { car_page %>% html_node("tr:nth-child(6) span") %>% html_text2() } # 提取车辆类别 get_car_category <- function(car_page) { car_page %>% html_node(".tw-col-span-6:nth-child(6)") %>% html_text2() } # 提取卖家信息 get_seller <- function(car_page) { car_page %>% html_node(".tw-my-2:nth-child(11) .tw-text-xl") %>% html_text2() } # -------------------------- # 主爬取循环 # -------------------------- for (page_result in seq(from=1, to=2710)) { link <- paste0("https://www.car.gr/classifieds/cars/?fs=1&condition=used&offer_type=sale&significant_damage=f&modified=15&pg=", page_result) # 请求列表页 page <- curl::curl(link) %>% read_html() Sys.sleep(sample(c(8,10,12), 1)) # 列表页间的延迟 # 提取列表页信息 Title <- page %>% html_nodes(".title") %>% html_text2() Price <- page %>% html_nodes(".price-fmt") %>% html_text2() car_links <- page %>% html_nodes(".row-anchor") %>% html_attr("href") %>% paste0("https://www.car.gr", .) # 批量获取所有车辆页面(每个链接仅请求一次,带延迟) car_pages <- lapply(car_links, function(link) { # 车辆页面请求前的随机延迟 Sys.sleep(sample(c(5,7,9), 1)) # 带容错的安全请求 possibly(get_car_page, otherwise = NULL)(link) }) # 过滤掉请求失败的页面 valid_indices <- !sapply(car_pages, is.null) car_pages_valid <- car_pages[valid_indices] Title_valid <- Title[valid_indices] Price_valid <- Price[valid_indices] # 基于已获取的页面提取各属性 color <- sapply(car_pages_valid, get_color) gearing_system <- sapply(car_pages_valid, get_gearing_sys) HP <- sapply(car_pages_valid, get_car_HP) CC <- sapply(car_pages_valid, get_car_CC) FUEL <- sapply(car_pages_valid, get_car_fuel) KM <- sapply(car_pages_valid, get_car_km) DATE <- sapply(car_pages_valid, get_car_date) CATEGORY <- sapply(car_pages_valid, get_car_category) SELLER <- sapply(car_pages_valid, get_seller) # 合并数据并追加到总数据框 current_page_data <- data.frame(Title_valid, Price_valid, color, gearing_system, HP, CC, FUEL, KM, DATE, CATEGORY, SELLER, stringsAsFactors = FALSE) CARS <- rbind(CARS, current_page_data) print(paste("Page", page_result, "of", "2710", "| Valid entries added:", nrow(current_page_data))) }
2. 优化延迟与重试机制
- 列表页间设置8-12秒随机延迟,车辆页面请求前设置5-9秒随机延迟,避免固定间隔被识别为爬虫
- 可添加重试逻辑,遇到429错误时自动延长延迟并重试:
# 带重试的页面请求函数 get_car_page_with_retry <- function(car_link) { max_attempts <- 3 for (attempt in 1:max_attempts) { tryCatch({ page <- read_html(car_link) return(page) }, error = function(e) { if (grepl("429", e$message)) { wait_time <- 10 * attempt # 重试延迟翻倍 Sys.sleep(wait_time) } else { stop(e) } }) } warning(paste("Failed to fetch", car_link, "after", max_attempts, "attempts")) return(NULL) } # 替换主循环中的get_car_page为该函数 car_pages <- lapply(car_links, function(link) { Sys.sleep(sample(c(5,7,9), 1)) get_car_page_with_retry(link) })
3. 修复原代码中的低级错误
- 原代码中
DATE<-sapply(car_links,get_car_km,simplify = TRUE)调用了错误的函数,应改为get_car_date SELLER<-SAPPLY(car_links,get_seller,simplify = TRUE)中sapply的大小写错误,应为小写get_seller函数中错误使用read_html(link)而非已获取的car_page,且返回变量未定义,已在上述代码中修复
4. 遵守网站爬虫规则
运行paths_allowed("https://www.car.gr/classifieds/cars/")确认网站是否允许爬取目标路径,避免违反robots.txt规则。
内容的提问来源于stack exchange,提问作者Spink
相关产品推荐
相关产品推荐

