You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免GitHub API速率限制导致Python程序崩溃?

问题描述

使用Python代码加载GitHub仓库内容时,触发github.GithubException.RateLimitExceededException,提示API速率超限。原本以为每个仓库仅调用一次API,但实际请求量超出GitHub每小时5000次的限制,需要可行的规避方案。

相关代码

from github import Github, Auth
import typing as T
from langchain.docstore.document import Document


def load_github_repos():
    def clone_github_repo(org_name, repo_name, files:T.Tuple=(".md", ".txt")) -> T.List[Document]:
        auth = Auth.Token(GIT_TOKEN)
        g = Github(auth=auth)
        repo = g.get_repo(f"{org_name}/{repo_name}")
        contents = repo.get_contents("")
        docs = []
        while contents:
            file_content = contents.pop(0)
            if file_content.type == "dir":
                contents.extend(repo.get_contents(file_content.path))
            else:
                if file_content.path.endswith(files):
                    docs.append(Document(page_content=file_content.decoded_content, metadata={"filename":file_content.name}))
        return docs
    repo_docs = []
    for repo in REPOS:
        repo_docs += clone_github_repo(repo_name=repo)

报错栈信息

Traceback (most recent call last):
  File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 174, in <module>
    update_vector_database()
  File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 115, in update_vector_database
    load_github_repos())
  File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 63, in load_github_repos
    repo_docs += clone_github_repo(repo_name=repo)
  File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 56, in clone_github_repo
    contents.extend(repo.get_contents(file_content.path))
  File "/Users/john.eastman/workspace/venv/lib/python3.9/site-packages/github/Repository.py", line 2107, in get_contents
    headers, data = self._requester.requestJsonAndCheck(
  File "/Users/john.eastman/workspace/venv/lib/python3.9/site-packages/github/Requester.py", line 442, in requestJsonAndCheck
    return self.__check(
  File "/Users/john.eastman/workspace/venv/lib/python3.9/site-packages/github/Requester.py", line 487, in __check
    raise self.__createException(status, responseHeaders, data)
github.GithubException.RateLimitExceededException: 403 {"message": "API rate limit exceeded for user ID 80288341.", "documentation_url": "https://docs.github.com/rest/overview/resources-in-the-rest-api#rate-limiting"}
解决方案

问题根源:代码中每遍历一个目录就调用一次repo.get_contents(),而非每个仓库仅调用一次。若仓库存在大量目录,会触发数十甚至数百次API请求,快速耗尽限额。以下是具体优化方案:

1. 复用Github客户端实例

当前代码每个仓库都创建新的Github实例,完全没必要。将客户端初始化移到外层,所有仓库共用一个实例:

def load_github_repos():
    # 把Github实例移到外层,所有仓库复用
    auth = Auth.Token(GIT_TOKEN)
    g = Github(auth=auth)

    def clone_github_repo(org_name, repo_name, files:T.Tuple=(".md", ".txt")) -> T.List[Document]:
        repo = g.get_repo(f"{org_name}/{repo_name}")
        # 后续代码保持不变

2. 一次性获取仓库所有文件,减少API调用

Github API支持通过recursive=True参数一次性拉取仓库所有文件,无需逐个目录请求。修改get_contents调用方式:

# 替换原有的contents初始化和循环逻辑
contents = repo.get_contents("", recursive=True)
docs = []
for file_content in contents:
    if file_content.type != "dir" and file_content.path.endswith(files):
        docs.append(Document(page_content=file_content.decoded_content, metadata={"filename":file_content.name}))

此修改后,单个仓库仅需一次API请求即可获取所有文件,大幅降低请求量。

3. 添加重试与速率控制

即使优化后,批量处理多仓库仍可能触发限额,需添加自动重试和延迟机制:

  • 安装tenacity库实现重试:pip install tenacity
  • 示例代码:
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
import github

def load_github_repos():
    auth = Auth.Token(GIT_TOKEN)
    g = Github(auth=auth)

    @retry(
        stop=stop_after_attempt(3),
        wait=wait_exponential(multiplier=1, min=2, max=10),
        retry=retry_if_exception_type(github.GithubException.RateLimitExceededException)
    )
    def clone_github_repo(org_name, repo_name, files:T.Tuple=(".md", ".txt")) -> T.List[Document]:
        repo = g.get_repo(f"{org_name}/{repo_name}")
        contents = repo.get_contents("", recursive=True)
        docs = []
        for file_content in contents:
            if file_content.type != "dir" and file_content.path.endswith(files):
                docs.append(Document(page_content=file_content.decoded_content, metadata={"filename":file_content.name}))
        return docs
    # 后续循环逻辑保持不变

4. 主动检查剩余API限额

在代码中主动查询当前剩余限额,当剩余请求不足时暂停等待:

import datetime
import time

def check_rate_limit(g):
    rate_limit = g.get_rate_limit()
    remaining = rate_limit.core.remaining
    reset_time = rate_limit.core.reset
    if remaining < 10:  # 剩余请求少于10个时触发等待
        wait_time = (reset_time - datetime.datetime.now(datetime.timezone.utc)).total_seconds() + 60  # 多等60秒确保限额重置
        time.sleep(wait_time)

# 在遍历仓库时调用检查函数
for repo in REPOS:
    check_rate_limit(g)
    repo_docs += clone_github_repo(repo_name=repo)

5. 改用GitHub GraphQL API(可选)

GraphQL API的限额规则更宽松(每小时5000次,但单请求可获取更多数据),适合批量操作。不过需要重构代码适配GraphQL语法,但能进一步减少请求次数。


内容的提问来源于stack exchange,提问作者Eastman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 21:07:21