You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用balldontlie API拉取全量stats数据报错求解决方案

解决方案

问题根源

  • 限速逻辑适配错误:官方标注的是每秒60次请求,而非每分钟60次。你当前的逻辑是短时间内连续发59次请求后休眠1分钟,虽然单秒请求量可能没超限,但很容易触发服务端的动态限流规则,且不同请求的响应耗时波动会导致速率控制完全失效。
  • 缺少异常与重试机制:没有判断HTTP响应状态码就直接解析JSON,当接口返回限流响应、网络波动、服务端临时报错时,直接调用response.json()会直接抛出异常中断爬取。
  • 没有速率冗余:刚好卡着上限发请求,服务端计数规则和本地计数规则存在差异时极易触限。

优化方案

  • 把请求速率控制在每秒50~55次,留出足够的冗余空间,避免触碰到服务端限流阈值
  • 增加HTTP状态码判断,遇到429限流响应时自动等待后重试,遇到5xx服务端错误、网络异常时也支持重试
  • 单次异常最多重试3次,避免死循环,重试失败的页码可以单独记录后续补爬
  • 不需要的meta字段直接丢弃,只保留需要的data内容

优化后代码

import requests
import json
import time

total_results = []
pages_to_read = 11000
# 速率控制:每秒最多55次请求,留出冗余避免触限
request_interval = 1 / 55
# 单页最多重试3次
max_retry = 3

last_request_time = 0

for page_num in range(1, pages_to_read + 1):
    retry_count = 0
    success = False
    while retry_count < max_retry and not success:
        # 控制请求间隔,保证速率稳定
        current_time = time.perf_counter()
        if current_time - last_request_time < request_interval:
            time.sleep(request_interval - (current_time - last_request_time))
        
        url = f"https://balldontlie.io/api/v1/stats?per_page=100&page={page_num}"
        print(f"正在读取第 {page_num} 页,重试次数:{retry_count}")
        
        try:
            response = requests.get(url, timeout=10)
            last_request_time = time.perf_counter()
            # 触发限流时等待10秒重试
            if response.status_code == 429:
                print("触发限流,等待10秒后重试")
                time.sleep(10)
                retry_count += 1
                continue
            # 服务端错误时重试
            if response.status_code >= 500:
                print(f"服务端错误,状态码:{response.status_code},重试")
                retry_count += 1
                time.sleep(2)
                continue
            # 正常响应解析数据
            data = response.json()
            if "data" in data:
                total_results.extend(data["data"])
            success = True
        except Exception as e:
            print(f"请求出错:{str(e)},重试")
            retry_count += 1
            time.sleep(2)
    
    if not success:
        print(f"第 {page_num} 页多次重试失败,已跳过,可后续补爬")

print(f"总共获取到 {len(total_results)} 条结果")

with open('test.json', 'w', encoding='utf-8') as d:
    json.dump(total_results, d, ensure_ascii=False, indent=4)

内容的提问来源于stack exchange,提问作者beds

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 06:36:02