如何强制GitHub API响应始终返回Content-Length请求头?
问题:GitHub API响应头不总是包含Content-Length,如何强制返回该字段?
在向GitHub API发起请求下载仓库压缩包时,发现响应头并不稳定包含Content-Length字段:
包含Content-Length的响应头示例
{ "Access-Control-Allow-Origin": "https://render.githubusercontent.com", "content-disposition": "attachment; filename=TheAmazingJeh-McSaveFiles-679514d.zip", "Content-Length": "230360822", "Content-Security-Policy": "default-src 'none'; style-src 'unsafe-inline'; sandbox", "Content-Type": "application/zip", "ETag": "W/\"43b5092d176697371a0d15d5f1e12f1cdac5d6c1ca0950d70cd4721b75105ead\"", "Strict-Transport-Security": "max-age=31536000", "Vary": "Authorization,Accept-Encoding,Origin", "X-Content-Type-Options": "nosniff", "X-Frame-Options": "deny", "X-XSS-Protection": "1; mode=block", "Date": "Wed, 03 May 2023 09:19:37 GMT", "X-GitHub-Request-Id": "9000:EF57:66D0E6:834C83:64522729" }
不包含Content-Length的响应头示例
{ "Access-Control-Allow-Origin": "https://render.githubusercontent.com", "content-disposition": "attachment; filename=TheAmazingJeh-McSaveFiles-679514d.zip", "Content-Security-Policy": "default-src 'none'; style-src 'unsafe-inline'; sandbox", "Content-Type": "application/zip", "ETag": "W/\"43b5092d176697371a0d15d5f1e12f1cdac5d6c1ca0950d70cd4721b75105ead\"", "Strict-Transport-Security": "max-age=31536000", "Vary": "Authorization,Accept-Encoding,Origin", "X-Content-Type-Options": "nosniff", "X-Frame-Options": "deny", "X-XSS-Protection": "1; mode=block", "Date": "Wed, 03 May 2023 09:17:49 GMT", "Transfer-Encoding": "chunked", "X-GitHub-Request-Id": "B1D9:06BC:5CB241:72F20E:645226BC" }
请问是否有办法强制GitHub API的响应始终包含Content-Length字段?
以下是当前使用的Python代码:
def download_zip(GITHUB_REPO, branch, zip_path, autoDownload=True): """Downloads a zip file from a Github repository autoDownload: If True, the file will be downloaded automatically. If False, the file will be opened in the browser, and the user will have to download it manually""" headers = { "Authorization" : f'{GITHUB_TOKEN} ghp_r5***', "Accept": '*.*', } g = Github(GITHUB_TOKEN) OWNER = g.get_user().name EXT = "zip" url = f'https://api.github.com/repos/{OWNER}/{GITHUB_REPO}/{EXT}ball/{branch}' if autoDownload: print(f"Requesting download for {c(branch, 'purple')} from {GITHUB_REPO}") for i in range(5): r = requests.get(url, headers=headers, stream=True) if r.headers.get("Content-Length") is not None: break print(f"Retrying... {i+1}/5", end="\r") if r.status_code == 200: print(f"Request successful, downloading file {c(branch, 'purple')}.{EXT}") with open(f'{zip_path}\\{branch}-download.{EXT}', 'wb') as f: total_length = r.headers.get('Content-Length') if total_length is None: # no content length header prRed("No content length header. File size will guessed.") total_length = len(r.content) dl = 0 downloaded_kb = 0 total_length = int(total_length) total_length_kb = int(round(total_length / 1024)) for data in r.iter_content(chunk_size=4096): # 4096 bytes downloaded_kb += int(round(len(data)/1024)) dl += len(data) f.write(data) done = int(50 * dl / total_length) prANSI(f" [{'=' * done}{' ' * (50-done)}] {downloaded_kb} kb /{total_length_kb} kb (~{int(round(total_length_kb/1024))} mb)", 222, end="\r") prGreen(f"Downloaded {branch} from {GITHUB_REPO} ({downloaded_kb} kb) ") else: print(r.text) prRed("Error downloading file") else: prPurple(f"Opening {url} in browser") webbrowser.open(url)
解答
无法强制GitHub API返回Content-Length
当服务器采用**分块编码(Transfer-Encoding: chunked)**传输数据时,不会返回Content-Length字段——因为数据是分块生成并发送的,总大小在传输前无法确定。GitHub API会根据请求负载、服务器状态等自动选择传输方式,没有公开的参数或设置可以强制它始终返回Content-Length。
优化代码的建议
当前代码依赖重试获取带Content-Length的响应,且无该字段时会调用len(r.content)(会把整个文件加载到内存,大文件易导致内存溢出),建议调整逻辑:
- 去掉不可靠的重试逻辑:重试无法保证一定能获取到带
Content-Length的响应,反而增加不必要的请求。 - 避免加载整个文件到内存:无
Content-Length时,直接实时显示已下载大小,而非预估总大小。 - 修正Authorization头格式:GitHub API的授权头格式应为
token <你的令牌>,原代码格式可能无效。 - 使用login而非name:用户的
name可能为空或不唯一,login是GitHub账户的唯一标识。
修改后的代码示例:
def download_zip(GITHUB_REPO, branch, zip_path, autoDownload=True): """Downloads a zip file from a Github repository autoDownload: If True, the file will be downloaded automatically. If False, the file will be opened in the browser, and the user will have to download it manually""" headers = { "Authorization": f"token {GITHUB_TOKEN}", "Accept": "*/*", } g = Github(GITHUB_TOKEN) OWNER = g.get_user().login EXT = "zip" url = f"https://api.github.com/repos/{OWNER}/{GITHUB_REPO}/{EXT}ball/{branch}" if autoDownload: print(f"Requesting download for {c(branch, 'purple')} from {GITHUB_REPO}") r = requests.get(url, headers=headers, stream=True) if r.status_code != 200: print(r.text) prRed("Error downloading file") return print(f"Request successful, downloading file {c(branch, 'purple')}.{EXT}") file_path = f"{zip_path}\\{branch}-download.{EXT}" with open(file_path, 'wb') as f: total_length = r.headers.get('Content-Length') downloaded_kb = 0 if total_length is not None: total_length = int(total_length) total_length_kb = int(round(total_length / 1024)) total_length_mb = int(round(total_length_kb / 1024)) for data in r.iter_content(chunk_size=4096): downloaded_kb += int(round(len(data)/1024)) f.write(data) done = int(50 * downloaded_kb * 1024 / total_length) prANSI(f" [{'=' * done}{' ' * (50-done)}] {downloaded_kb} kb / {total_length_kb} kb (~{total_length_mb} mb)", 222, end="\r") else: prRed("No content length header. Showing downloaded size only.") for data in r.iter_content(chunk_size=4096): downloaded_kb += int(round(len(data)/1024)) f.write(data) prANSI(f" Downloaded: {downloaded_kb} kb (~{int(round(downloaded_kb/1024))} mb)", 222, end="\r") prGreen(f"Downloaded {branch} from {GITHUB_REPO} ({downloaded_kb} kb) ") else: prPurple(f"Opening {url} in browser") webbrowser.open(url)
内容的提问来源于stack exchange,提问作者TheAmazingJeh
相关产品推荐
相关产品推荐

