You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何强制GitHub API响应始终返回Content-Length请求头?

问题:GitHub API响应头不总是包含Content-Length,如何强制返回该字段?

在向GitHub API发起请求下载仓库压缩包时,发现响应头并不稳定包含Content-Length字段:

包含Content-Length的响应头示例

{
  "Access-Control-Allow-Origin": "https://render.githubusercontent.com",
  "content-disposition": "attachment; filename=TheAmazingJeh-McSaveFiles-679514d.zip",
  "Content-Length": "230360822",
  "Content-Security-Policy": "default-src 'none'; style-src 'unsafe-inline'; sandbox",
  "Content-Type": "application/zip",
  "ETag": "W/\"43b5092d176697371a0d15d5f1e12f1cdac5d6c1ca0950d70cd4721b75105ead\"",
  "Strict-Transport-Security": "max-age=31536000",
  "Vary": "Authorization,Accept-Encoding,Origin",
  "X-Content-Type-Options": "nosniff",
  "X-Frame-Options": "deny",
  "X-XSS-Protection": "1; mode=block",
  "Date": "Wed, 03 May 2023 09:19:37 GMT",
  "X-GitHub-Request-Id": "9000:EF57:66D0E6:834C83:64522729"
}

不包含Content-Length的响应头示例

{
  "Access-Control-Allow-Origin": "https://render.githubusercontent.com",
  "content-disposition": "attachment; filename=TheAmazingJeh-McSaveFiles-679514d.zip",
  "Content-Security-Policy": "default-src 'none'; style-src 'unsafe-inline'; sandbox",
  "Content-Type": "application/zip",
  "ETag": "W/\"43b5092d176697371a0d15d5f1e12f1cdac5d6c1ca0950d70cd4721b75105ead\"",
  "Strict-Transport-Security": "max-age=31536000",
  "Vary": "Authorization,Accept-Encoding,Origin",
  "X-Content-Type-Options": "nosniff",
  "X-Frame-Options": "deny",
  "X-XSS-Protection": "1; mode=block",
  "Date": "Wed, 03 May 2023 09:17:49 GMT",
  "Transfer-Encoding": "chunked",
  "X-GitHub-Request-Id": "B1D9:06BC:5CB241:72F20E:645226BC"
}

请问是否有办法强制GitHub API的响应始终包含Content-Length字段?

以下是当前使用的Python代码:

def download_zip(GITHUB_REPO, branch, zip_path, autoDownload=True):
    """Downloads a zip file from a Github repository
    
    autoDownload: If True, the file will be downloaded automatically. If False, the file will be opened in the browser, and the user will have to download it manually"""
    headers = {
        "Authorization" : f'{GITHUB_TOKEN} ghp_r5***',
        "Accept": '*.*',

    }

    g = Github(GITHUB_TOKEN)
    OWNER = g.get_user().name
    EXT = "zip"
    url = f'https://api.github.com/repos/{OWNER}/{GITHUB_REPO}/{EXT}ball/{branch}'

    if autoDownload:
        print(f"Requesting download for {c(branch, 'purple')} from {GITHUB_REPO}")
        
        for i in range(5):
            r = requests.get(url, headers=headers, stream=True)
            if r.headers.get("Content-Length") is not None: break
            print(f"Retrying... {i+1}/5", end="\r")

        if r.status_code == 200:
            print(f"Request successful, downloading file {c(branch, 'purple')}.{EXT}")
            with open(f'{zip_path}\\{branch}-download.{EXT}', 'wb') as f:
                total_length = r.headers.get('Content-Length')

                if total_length is None: # no content length header
                    prRed("No content length header. File size will guessed.")
                    total_length = len(r.content)
                dl = 0
                downloaded_kb = 0
                total_length = int(total_length)
                total_length_kb = int(round(total_length / 1024))
                for data in r.iter_content(chunk_size=4096): # 4096 bytes
                    downloaded_kb += int(round(len(data)/1024))
                    dl += len(data)
                    f.write(data)
                    done = int(50 * dl / total_length) 
                    prANSI(f"  [{'=' * done}{' ' * (50-done)}] {downloaded_kb} kb /{total_length_kb} kb (~{int(round(total_length_kb/1024))} mb)", 222, end="\r")
            prGreen(f"Downloaded {branch} from {GITHUB_REPO} ({downloaded_kb} kb)                                         ")
        else:
            print(r.text)   
            prRed("Error downloading file")
    else:
        prPurple(f"Opening {url} in browser")
        webbrowser.open(url)

解答

无法强制GitHub API返回Content-Length

当服务器采用**分块编码(Transfer-Encoding: chunked)**传输数据时,不会返回Content-Length字段——因为数据是分块生成并发送的,总大小在传输前无法确定。GitHub API会根据请求负载、服务器状态等自动选择传输方式,没有公开的参数或设置可以强制它始终返回Content-Length。

优化代码的建议

当前代码依赖重试获取带Content-Length的响应,且无该字段时会调用len(r.content)(会把整个文件加载到内存,大文件易导致内存溢出),建议调整逻辑:

  1. 去掉不可靠的重试逻辑:重试无法保证一定能获取到带Content-Length的响应,反而增加不必要的请求。
  2. 避免加载整个文件到内存:无Content-Length时,直接实时显示已下载大小,而非预估总大小。
  3. 修正Authorization头格式:GitHub API的授权头格式应为token <你的令牌>,原代码格式可能无效。
  4. 使用login而非name:用户的name可能为空或不唯一,login是GitHub账户的唯一标识。

修改后的代码示例:

def download_zip(GITHUB_REPO, branch, zip_path, autoDownload=True):
    """Downloads a zip file from a Github repository
    
    autoDownload: If True, the file will be downloaded automatically. If False, the file will be opened in the browser, and the user will have to download it manually"""
    headers = {
        "Authorization": f"token {GITHUB_TOKEN}",
        "Accept": "*/*",
    }

    g = Github(GITHUB_TOKEN)
    OWNER = g.get_user().login
    EXT = "zip"
    url = f"https://api.github.com/repos/{OWNER}/{GITHUB_REPO}/{EXT}ball/{branch}"

    if autoDownload:
        print(f"Requesting download for {c(branch, 'purple')} from {GITHUB_REPO}")
        
        r = requests.get(url, headers=headers, stream=True)
        if r.status_code != 200:
            print(r.text)   
            prRed("Error downloading file")
            return

        print(f"Request successful, downloading file {c(branch, 'purple')}.{EXT}")
        file_path = f"{zip_path}\\{branch}-download.{EXT}"
        with open(file_path, 'wb') as f:
            total_length = r.headers.get('Content-Length')
            downloaded_kb = 0

            if total_length is not None:
                total_length = int(total_length)
                total_length_kb = int(round(total_length / 1024))
                total_length_mb = int(round(total_length_kb / 1024))
                for data in r.iter_content(chunk_size=4096):
                    downloaded_kb += int(round(len(data)/1024))
                    f.write(data)
                    done = int(50 * downloaded_kb * 1024 / total_length)
                    prANSI(f"  [{'=' * done}{' ' * (50-done)}] {downloaded_kb} kb / {total_length_kb} kb (~{total_length_mb} mb)", 222, end="\r")
            else:
                prRed("No content length header. Showing downloaded size only.")
                for data in r.iter_content(chunk_size=4096):
                    downloaded_kb += int(round(len(data)/1024))
                    f.write(data)
                    prANSI(f"  Downloaded: {downloaded_kb} kb (~{int(round(downloaded_kb/1024))} mb)", 222, end="\r")
        
        prGreen(f"Downloaded {branch} from {GITHUB_REPO} ({downloaded_kb} kb)                                         ")
    else:
        prPurple(f"Opening {url} in browser")
        webbrowser.open(url)

内容的提问来源于stack exchange,提问作者TheAmazingJeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 09:37:08