如何避免GitHub API速率限制导致Python程序崩溃?
问题描述
使用Python代码加载GitHub仓库内容时,触发github.GithubException.RateLimitExceededException,提示API速率超限。原本以为每个仓库仅调用一次API,但实际请求量超出GitHub每小时5000次的限制,需要可行的规避方案。
相关代码
from github import Github, Auth import typing as T from langchain.docstore.document import Document def load_github_repos(): def clone_github_repo(org_name, repo_name, files:T.Tuple=(".md", ".txt")) -> T.List[Document]: auth = Auth.Token(GIT_TOKEN) g = Github(auth=auth) repo = g.get_repo(f"{org_name}/{repo_name}") contents = repo.get_contents("") docs = [] while contents: file_content = contents.pop(0) if file_content.type == "dir": contents.extend(repo.get_contents(file_content.path)) else: if file_content.path.endswith(files): docs.append(Document(page_content=file_content.decoded_content, metadata={"filename":file_content.name})) return docs repo_docs = [] for repo in REPOS: repo_docs += clone_github_repo(repo_name=repo)
报错栈信息
Traceback (most recent call last): File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 174, in <module> update_vector_database() File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 115, in update_vector_database load_github_repos()) File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 63, in load_github_repos repo_docs += clone_github_repo(repo_name=repo) File "/Users/john.eastman/workspace/rfp-monster/data_scripts/update_db.py", line 56, in clone_github_repo contents.extend(repo.get_contents(file_content.path)) File "/Users/john.eastman/workspace/venv/lib/python3.9/site-packages/github/Repository.py", line 2107, in get_contents headers, data = self._requester.requestJsonAndCheck( File "/Users/john.eastman/workspace/venv/lib/python3.9/site-packages/github/Requester.py", line 442, in requestJsonAndCheck return self.__check( File "/Users/john.eastman/workspace/venv/lib/python3.9/site-packages/github/Requester.py", line 487, in __check raise self.__createException(status, responseHeaders, data) github.GithubException.RateLimitExceededException: 403 {"message": "API rate limit exceeded for user ID 80288341.", "documentation_url": "https://docs.github.com/rest/overview/resources-in-the-rest-api#rate-limiting"}
解决方案
问题根源:代码中每遍历一个目录就调用一次repo.get_contents(),而非每个仓库仅调用一次。若仓库存在大量目录,会触发数十甚至数百次API请求,快速耗尽限额。以下是具体优化方案:
1. 复用Github客户端实例
当前代码每个仓库都创建新的Github实例,完全没必要。将客户端初始化移到外层,所有仓库共用一个实例:
def load_github_repos(): # 把Github实例移到外层,所有仓库复用 auth = Auth.Token(GIT_TOKEN) g = Github(auth=auth) def clone_github_repo(org_name, repo_name, files:T.Tuple=(".md", ".txt")) -> T.List[Document]: repo = g.get_repo(f"{org_name}/{repo_name}") # 后续代码保持不变
2. 一次性获取仓库所有文件,减少API调用
Github API支持通过recursive=True参数一次性拉取仓库所有文件,无需逐个目录请求。修改get_contents调用方式:
# 替换原有的contents初始化和循环逻辑 contents = repo.get_contents("", recursive=True) docs = [] for file_content in contents: if file_content.type != "dir" and file_content.path.endswith(files): docs.append(Document(page_content=file_content.decoded_content, metadata={"filename":file_content.name}))
此修改后,单个仓库仅需一次API请求即可获取所有文件,大幅降低请求量。
3. 添加重试与速率控制
即使优化后,批量处理多仓库仍可能触发限额,需添加自动重试和延迟机制:
- 安装
tenacity库实现重试:pip install tenacity - 示例代码:
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type import github def load_github_repos(): auth = Auth.Token(GIT_TOKEN) g = Github(auth=auth) @retry( stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10), retry=retry_if_exception_type(github.GithubException.RateLimitExceededException) ) def clone_github_repo(org_name, repo_name, files:T.Tuple=(".md", ".txt")) -> T.List[Document]: repo = g.get_repo(f"{org_name}/{repo_name}") contents = repo.get_contents("", recursive=True) docs = [] for file_content in contents: if file_content.type != "dir" and file_content.path.endswith(files): docs.append(Document(page_content=file_content.decoded_content, metadata={"filename":file_content.name})) return docs # 后续循环逻辑保持不变
4. 主动检查剩余API限额
在代码中主动查询当前剩余限额,当剩余请求不足时暂停等待:
import datetime import time def check_rate_limit(g): rate_limit = g.get_rate_limit() remaining = rate_limit.core.remaining reset_time = rate_limit.core.reset if remaining < 10: # 剩余请求少于10个时触发等待 wait_time = (reset_time - datetime.datetime.now(datetime.timezone.utc)).total_seconds() + 60 # 多等60秒确保限额重置 time.sleep(wait_time) # 在遍历仓库时调用检查函数 for repo in REPOS: check_rate_limit(g) repo_docs += clone_github_repo(repo_name=repo)
5. 改用GitHub GraphQL API(可选)
GraphQL API的限额规则更宽松(每小时5000次,但单请求可获取更多数据),适合批量操作。不过需要重构代码适配GraphQL语法,但能进一步减少请求次数。
内容的提问来源于stack exchange,提问作者Eastman
相关产品推荐
相关产品推荐

