动态网页爬取优化求助:rvest+chromote爬取Hoopshype薪资超时问题
问题描述
尝试爬取Hoopshype网站1990-2024年的球员薪资数据,页面URL结构统一(如1990年页面地址:https://www.hoopshype.com/salaries/players/?season=1990),但每个年份页面需要点击底部按钮加载后续表格。使用rvest和chromote的read_html_live结合循环编写了代码,但运行耗时极长且频繁超时,期望得到包含rank、player、salary三列的DataFrame,想优化爬取速度,询问是否可实现向量化及改进建议。
原代码如下:
library(rvest) library(chromote) start_time <- Sys.time() season_range <- seq(1990, 2024, by =001) year <- numeric() for(season in season_range){ year <- c(year, paste(season)) site <- paste('https://hoopshype.com/salaries/players/?season=year', year, sep = "") session <- read_html_live(site) all_data <- list() page_count <- 1 repeat { # 2. Extract table data from the current view current_table <- session %>% html_element("table") %>% html_table() all_data[[page_count]] <- current_table # 3. Check for the 'Next' button next_button <- session %>% html_element("#__next > div > div.cHtQSi__cHtQSi > div > div > div.cVmJc5__cVmJc5 > div:nth-child(2) > div.nsCjpX__nsCjpX.EJGS6o__EJGS6o > div.PndCGL__PndCGL.K0B27p__K0B27p > button.hd3Vfp__hd3Vfp._3JhbLM__3JhbLM > span") # 4. Exit if button is missing or disabled if (is.na(next_button)) break # 5. Click and wait for the new data to load session$click("#__next > div > div.cHtQSi__cHtQSi > div > div > div.cVmJc5__cVmJc5 > div:nth-child(2) > div.nsCjpX__nsCjpX.EJGS6o__EJGS6o > div.PndCGL__PndCGL.K0B27p__K0B27p > button.hd3Vfp__hd3Vfp._3JhbLM__3JhbLM > span") Sys.sleep(2) # Give the page time to update page_count <- page_count + 1 } } end_time <- Sys.time() print(end_time - start_time) # Combine all pages into one data frame final_df <- do.call(rbind, all_data)
问题分析
原代码存在几个核心问题导致效率低下和超时:
- URL拼接错误:将
year作为字符串拼接进URL,实际应该用当前循环的season变量生成正确地址 - 资源泄漏:每个年份新建Chromote会话但未关闭,导致浏览器进程堆积,占用大量资源
- 固定等待浪费时间:
Sys.sleep(2)不管页面实际加载状态,强制等待,增加不必要耗时 - 选择器脆弱:使用冗长的嵌套CSS选择器,页面结构微小变化就会导致选择失效
- 单线程串行处理:逐个年份爬取,未利用多核资源
优化方案
1. 修复基础错误
首先修正URL生成逻辑,确保每个年份的地址正确:
site <- paste0('https://hoopshype.com/salaries/players/?season=', season)
2. 优化会话管理
每次处理完一个年份后关闭Chromote会话,避免资源堆积:
# 在函数内添加退出时自动关闭会话 on.exit(session$close())
3. 替换固定等待为智能等待
使用Chromote的wait_for方法,等待页面元素加载完成后再继续,避免无效等待:
# 点击后等待表格行数增加,确认新数据加载完成 prev_rows <- nrow(current_table) session$wait_for(paste0("document.querySelector('table').rows.length > ", prev_rows), timeout = 10000)
4. 简化选择器提升稳定性
用XPath匹配按钮文本,替代脆弱的长CSS选择器:
# 查找Next按钮 next_button <- session %>% html_element(xpath = "//button[contains(text(), 'Next')]") # 点击按钮 session$click(xpath = "//button[contains(text(), 'Next')]")
5. 并行处理提升爬取速度
年份之间是独立任务,可通过并行库furrr实现多年份同时爬取,大幅缩短总耗时:
library(rvest) library(chromote) library(furrr) library(dplyr) # 设置并行工作数(建议2-4,避免触发反爬) plan(multisession, workers = 3) # 定义单年份爬取函数 scrape_season <- function(season) { # 初始化会话 session <- read_html_live(paste0('https://hoopshype.com/salaries/players/?season=', season)) # 退出时自动关闭会话 on.exit(session$close()) all_data <- list() page_count <- 1 repeat { # 提取表格并只保留需要的列,添加年份标识 current_table <- session %>% html_element("table") %>% html_table() %>% select(Rank, Player, Salary) %>% # 根据实际页面列名调整 mutate(Season = season) all_data[[page_count]] <- current_table # 检查Next按钮是否存在且可用 next_button <- session %>% html_element(xpath = "//button[contains(text(), 'Next')]") if (is.na(next_button) || session$evaluate_js("arguments[0].disabled", next_button)) { break } # 点击按钮并等待数据加载 session$click(xpath = "//button[contains(text(), 'Next')]") prev_rows <- nrow(current_table) session$wait_for(paste0("document.querySelector('table').rows.length > ", prev_rows), timeout = 10000) page_count <- page_count + 1 } # 合并单年份所有页面数据 do.call(rbind, all_data) } # 并行爬取所有年份 season_range <- 1990:2024 start_time <- Sys.time() final_df <- future_map_dfr(season_range, scrape_season) end_time <- Sys.time() print(end_time - start_time)
6. 反爬注意事项
- 控制并行数,不要超过4,避免被网站限制
- 可添加随机等待时间,比如
Sys.sleep(runif(1, 0.5, 1.5)),模拟人工操作 - 若遇到频繁限制,可尝试配置Chromote的User-Agent,模拟真实浏览器
关于向量化的说明
动态页面爬取依赖浏览器交互(点击、等待加载),这类操作本质是串行的,无法直接实现向量化。但通过并行处理可以将多个年份的爬取任务同时执行,达到类似"向量化"的效率提升效果,这是当前场景下最优的提速方案。
内容的提问来源于stack exchange,提问作者jvalenti
相关产品推荐
相关产品推荐

