You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何我的GitHub代码搜索触发二次速率限制?求优化方案

解决GitHub搜索API二次速率限制问题

问题背景

我正在搜索GitHub中包含字符串“torch.”的Python文件。由于GitHub搜索API仅返回前100条结果,我按照相关建议按文件大小分区间(0到500000,步长250)分页查询,但持续触发二次速率限制。已查阅官方集成最佳实践仍未解决,以下是我的代码:

import os
import requests
import httplink
import time

# This for loop searches for code based on files sizes from 0 to 500000 containing the string "torch"
for i in range(0,500000,250):
  print("i = ",i," i + 250 = ", i+250)
  url = "https://api.github.com/search/code?q=torch +in:file + language:python+size:"+str(i)+".."+str(i+250)+"&page=1&per_page=10" 

  headers = {"Authorization": f'Token xxxxxxxxxxxxxxx'} ## Please put your token over here

  # Backoff when secondary rate limit is reached
  backoff = 256

  total = 0
  cond = True

  # This while loop goes over all pages of results => Pagination
  while cond==True:
    try:
      

          time.sleep(2)
          res = requests.request("GET", url, headers=headers)
          res.raise_for_status()
          link = httplink.parse_link_header(res.headers["link"])

          data = res.json()
          for i, item in enumerate(data["items"], start=total):
              print(f'[{i}] {item["html_url"]}')

          if "next" not in link:
              break

          total += len(data["items"])

          url = link["next"].target

    # Except case to catch when secondary rate limit has been reached and prevent the computation from stopping
    except requests.exceptions.HTTPError as err:
        print("err = ", err)
        print("err.response.text = ", err.response.text)
        # backoff **= 2
        print("backoff = ", backoff)
        time.sleep(backoff)
    # Except case to catch when the given file size provides no results
    except KeyError as error:
      print("err = ", error)

      # Set cond to False to stop the while loop
      cond = False
      continue

可行优化方案

1. 动态遵循速率限制头部信息

不要固定sleep(2),而是读取GitHub API返回的X-RateLimit-Remaining和X-RateLimit-Reset头部字段:

  • 当剩余请求数低于10时,计算距离重置时间的剩余秒数,休眠对应时长再继续请求
  • 每次请求后根据头部数据调整策略,避免盲目休眠

2. 优化指数退避+随机抖动

当前固定backoff值的方式不够灵活,改用指数退避策略:

  • 初始退避时间设为1秒,触发限制后翻倍,最大不超过1024秒
  • 加入随机抖动(比如±2秒),避免多线程/批量请求扎堆触发限制

3. 减少无效请求次数

  • 将per_page调整为API允许的最大值30,减少分页请求的总次数
  • 增大文件大小区间步长(比如从250改为1000),跳过大量无匹配结果的空区间,减少整体请求批次

4. 修复代码逻辑缺陷

  • 外层循环变量i和内层枚举的i冲突,导致外层循环逻辑出错,把内层变量名改为idx
  • 处理link头部不存在的情况:当结果只有一页时,res.headers没有link字段,直接解析会报错,需要先判断
  • 触发HTTPError后,应该重新发起当前请求,而不是直接休眠后继续,避免跳过当前页数据

5. 可选:改用GitHub App身份验证

个人令牌的速率限制配额有限,如果大规模搜索需求,可改用GitHub App身份验证,能获得更高的请求配额。

改进后的代码示例

import os
import requests
import httplink
import time
import random

# 搜索含"torch."的Python文件,按文件大小区间分页
for size_start in range(0, 500000, 1000):
    size_end = size_start + 1000
    print(f"size range: {size_start}..{size_end}")
    url = f"https://api.github.com/search/code?q=torch.+in:file+language:python+size:{size_start}..{size_end}&page=1&per_page=30" 

    headers = {"Authorization": f'Token xxxxxxxxxxxxxxx'}
    backoff = 1
    total = 0
    cond = True

    while cond:
        try:
            res = requests.get(url, headers=headers)
            
            # 处理速率限制头部
            remaining = int(res.headers.get('X-RateLimit-Remaining', 0))
            reset_time = int(res.headers.get('X-RateLimit-Reset', time.time()))
            current_time = time.time()
            
            if remaining < 10:
                sleep_sec = reset_time - current_time + 1
                if sleep_sec > 0:
                    print(f"接近速率限制,休眠{sleep_sec}秒")
                    time.sleep(sleep_sec)
            
            res.raise_for_status()
            
            # 处理link头部
            link = None
            if "link" in res.headers:
                link = httplink.parse_link_header(res.headers["link"])
            
            data = res.json()
            if "items" not in data:
                cond = False
                continue
                
            for idx, item in enumerate(data["items"], start=total):
                print(f'[{idx}] {item["html_url"]}')
            
            total += len(data["items"])
            
            # 检查是否有下一页
            if not link or "next" not in link:
                break
            
            url = link["next"].target
            # 基础休眠+随机抖动,避免请求过于频繁
            time.sleep(1 + random.uniform(0, 0.5))
            # 重置退避时间
            backoff = 1

        except requests.exceptions.HTTPError as err:
            print(f"错误: {err}")
            print(f"响应内容: {err.response.text}")
            
            # 指数退避+随机抖动
            backoff = min(backoff * 2, 1024)
            sleep_time = backoff + random.uniform(0, 2)
            print(f"触发限制,休眠{sleep_time:.2f}秒")
            time.sleep(sleep_time)
            
            # 重新发起当前请求,不推进到下一页
            continue
            
        except KeyError as error:
            print(f"无匹配结果: {error}")
            cond = False
            continue

内容的提问来源于stack exchange,提问作者desert_ranger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 22:55:37