You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

动态网页爬取优化求助:rvest+chromote爬取Hoopshype薪资超时问题

问题描述

尝试爬取Hoopshype网站1990-2024年的球员薪资数据,页面URL结构统一(如1990年页面地址:https://www.hoopshype.com/salaries/players/?season=1990),但每个年份页面需要点击底部按钮加载后续表格。使用rvest和chromote的read_html_live结合循环编写了代码,但运行耗时极长且频繁超时,期望得到包含rank、player、salary三列的DataFrame,想优化爬取速度,询问是否可实现向量化及改进建议。

原代码如下:

library(rvest)
library(chromote)

start_time <- Sys.time()

season_range <- seq(1990, 2024, by =001)
year <- numeric()

for(season in season_range){
  year <- c(year, paste(season))

site <- paste('https://hoopshype.com/salaries/players/?season=year',
               year, sep = "")

session <- read_html_live(site)

all_data <- list()
page_count <- 1

repeat {
  # 2. Extract table data from the current view
  current_table <- session %>% 
    html_element("table") %>% 
    html_table()
  
  all_data[[page_count]] <- current_table
  
  # 3. Check for the 'Next' button
  next_button <- session %>% html_element("#__next > div > div.cHtQSi__cHtQSi > div > div > div.cVmJc5__cVmJc5 > div:nth-child(2) > div.nsCjpX__nsCjpX.EJGS6o__EJGS6o > div.PndCGL__PndCGL.K0B27p__K0B27p > button.hd3Vfp__hd3Vfp._3JhbLM__3JhbLM > span")
  
  # 4. Exit if button is missing or disabled
  if (is.na(next_button)) break
  
  # 5. Click and wait for the new data to load
  session$click("#__next > div > div.cHtQSi__cHtQSi > div > div > div.cVmJc5__cVmJc5 > div:nth-child(2) > div.nsCjpX__nsCjpX.EJGS6o__EJGS6o > div.PndCGL__PndCGL.K0B27p__K0B27p > button.hd3Vfp__hd3Vfp._3JhbLM__3JhbLM > span")
  Sys.sleep(2)  # Give the page time to update
  
  page_count <- page_count + 1
  }
}

end_time <- Sys.time()
print(end_time - start_time)

# Combine all pages into one data frame
final_df <- do.call(rbind, all_data)
问题分析

原代码存在几个核心问题导致效率低下和超时:

  • URL拼接错误:将year作为字符串拼接进URL,实际应该用当前循环的season变量生成正确地址
  • 资源泄漏:每个年份新建Chromote会话但未关闭,导致浏览器进程堆积,占用大量资源
  • 固定等待浪费时间:Sys.sleep(2)不管页面实际加载状态,强制等待,增加不必要耗时
  • 选择器脆弱:使用冗长的嵌套CSS选择器,页面结构微小变化就会导致选择失效
  • 单线程串行处理:逐个年份爬取,未利用多核资源
优化方案

1. 修复基础错误

首先修正URL生成逻辑,确保每个年份的地址正确:

site <- paste0('https://hoopshype.com/salaries/players/?season=', season)

2. 优化会话管理

每次处理完一个年份后关闭Chromote会话,避免资源堆积:

# 在函数内添加退出时自动关闭会话
on.exit(session$close())

3. 替换固定等待为智能等待

使用Chromote的wait_for方法,等待页面元素加载完成后再继续,避免无效等待:

# 点击后等待表格行数增加,确认新数据加载完成
prev_rows <- nrow(current_table)
session$wait_for(paste0("document.querySelector('table').rows.length > ", prev_rows), timeout = 10000)

4. 简化选择器提升稳定性

用XPath匹配按钮文本,替代脆弱的长CSS选择器:

# 查找Next按钮
next_button <- session %>% html_element(xpath = "//button[contains(text(), 'Next')]")
# 点击按钮
session$click(xpath = "//button[contains(text(), 'Next')]")

5. 并行处理提升爬取速度

年份之间是独立任务,可通过并行库furrr实现多年份同时爬取,大幅缩短总耗时:

library(rvest)
library(chromote)
library(furrr)
library(dplyr)

# 设置并行工作数(建议2-4,避免触发反爬)
plan(multisession, workers = 3)

# 定义单年份爬取函数
scrape_season <- function(season) {
  # 初始化会话
  session <- read_html_live(paste0('https://hoopshype.com/salaries/players/?season=', season))
  # 退出时自动关闭会话
  on.exit(session$close())
  
  all_data <- list()
  page_count <- 1
  
  repeat {
    # 提取表格并只保留需要的列,添加年份标识
    current_table <- session %>% 
      html_element("table") %>% 
      html_table() %>% 
      select(Rank, Player, Salary) %>% # 根据实际页面列名调整
      mutate(Season = season)
    
    all_data[[page_count]] <- current_table
    
    # 检查Next按钮是否存在且可用
    next_button <- session %>% html_element(xpath = "//button[contains(text(), 'Next')]")
    if (is.na(next_button) || session$evaluate_js("arguments[0].disabled", next_button)) {
      break
    }
    
    # 点击按钮并等待数据加载
    session$click(xpath = "//button[contains(text(), 'Next')]")
    prev_rows <- nrow(current_table)
    session$wait_for(paste0("document.querySelector('table').rows.length > ", prev_rows), timeout = 10000)
    
    page_count <- page_count + 1
  }
  
  # 合并单年份所有页面数据
  do.call(rbind, all_data)
}

# 并行爬取所有年份
season_range <- 1990:2024
start_time <- Sys.time()
final_df <- future_map_dfr(season_range, scrape_season)
end_time <- Sys.time()
print(end_time - start_time)

6. 反爬注意事项

  • 控制并行数,不要超过4,避免被网站限制
  • 可添加随机等待时间,比如Sys.sleep(runif(1, 0.5, 1.5)),模拟人工操作
  • 若遇到频繁限制,可尝试配置Chromote的User-Agent,模拟真实浏览器
关于向量化的说明

动态页面爬取依赖浏览器交互(点击、等待加载),这类操作本质是串行的,无法直接实现向量化。但通过并行处理可以将多个年份的爬取任务同时执行,达到类似"向量化"的效率提升效果,这是当前场景下最优的提速方案。

内容的提问来源于stack exchange,提问作者jvalenti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 14:54:53