You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用代理轮换爬取Basketball Reference仍遇429错误求解决方案

问题分析与修复方案

1. 修复代理代码的核心错误

你的代理版本代码存在逻辑错误:pd.read_csv用于解析CSV格式数据,但目标页面是HTML结构,应该沿用原逻辑的pd.read_html解析表格。调整后的代码示例:

import requests
import pandas as pd
from io import StringIO

link = 'https://www.basketball-reference.com/teams/NOP/2023.html#advanced'
# temp为你的代理地址,格式如 'http://123.45.67.89:8080'
proxies = {'http': temp, 'https': temp}
# 模拟浏览器请求头,避免被识别为爬虫
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

try:
    response = requests.get(link, proxies=proxies, headers=headers, timeout=5)
    response.raise_for_status()  # 主动抛出HTTP请求错误
    dfs = pd.read_html(StringIO(response.text))
    stats = dfs[1].dropna()
    d = stats.to_dict('records')
except Exception as e:
    print(f"请求或解析失败: {str(e)}")

2. 解决429错误的关键优化

  • 代理有效性校验:提前测试代理池中的代理能否正常访问目标网站,过滤掉超时、被封禁的无效代理
  • 请求频率控制:每次请求后添加随机延迟(2-5秒),避免短时间内发起大量请求
  • 重试与代理轮换:遇到429或请求失败时,自动切换代理并重试(可借助tenacity库实现重试逻辑)
  • 完善请求头:除User-Agent外,可添加Accept、Accept-Language等字段,进一步模拟浏览器行为
  • 使用会话保持:用requests.Session()复用连接,减少重复建立TCP连接的开销,降低被识别的概率

3. 额外注意事项

Basketball Reference有严格的反爬机制,即使使用代理,也不要一次性爬取大量页面;若数据需求较大,优先查看网站是否提供官方API或数据导出入口。

内容的提问来源于stack exchange,提问作者apersoninneed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 22:57:18