You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python请求库下载Zip文件仅获部分无效文件的问题求助

解决Zip文件下载不完整/IncompleteRead错误的方案

核心问题排查与解决步骤:

  1. 开启响应内容自动解码
    requests的r.raw默认不会解码分块传输编码的内容,导致下载文件混入分块标记,直接损坏Zip结构。修改代码添加解码配置:
try:
    r = requests.get(BASEURL, stream=True)
    r.raw.decode_content = True  # 关键:让raw自动处理分块编码
    with open(localfiletitle, 'wb') as fd:
        shutil.copyfileobj(r.raw, fd)
except requests.exceptions.RequestException as e:
    send_notice_mail("Error downloading the file:", e)
    return False
  1. 验证请求状态码
    先确认服务器返回成功响应,避免下载错误页面而非目标文件:
try:
    r = requests.get(BASEURL, stream=True)
    r.raise_for_status()  # 主动抛出4xx/5xx状态码的异常
    r.raw.decode_content = True
    with open(localfiletitle, 'wb') as fd:
        shutil.copyfileobj(r.raw, fd)
except requests.exceptions.RequestException as e:
    send_notice_mail("Error downloading the file:", e)
    return False
  1. 改用iter_content分块读取
    替代直接操作r.raw,使用requests提供的更稳定的分块读取方法:
try:
    r = requests.get(BASEURL, stream=True)
    r.raise_for_status()
    with open(localfiletitle, 'wb') as fd:
        # 按8KB块读取,可根据网络情况调整大小
        for chunk in r.iter_content(chunk_size=8192):
            if chunk:  # 过滤空块
                fd.write(chunk)
except requests.exceptions.RequestException as e:
    send_notice_mail("Error downloading the file:", e)
    return False
  1. 添加浏览器用户代理
    部分服务器会拦截默认requests代理,返回不完整内容,模拟浏览器请求:
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
try:
    r = requests.get(BASEURL, stream=True, headers=headers)
    r.raise_for_status()
    r.raw.decode_content = True
    with open(localfiletitle, 'wb') as fd:
        shutil.copyfileobj(r.raw, fd)
except requests.exceptions.RequestException as e:
    send_notice_mail("Error downloading the file:", e)
    return False
  1. urllib的IncompleteRead异常处理
    如果要继续使用urllib,可捕获异常并重试:
import urllib.request
from http.client import IncompleteRead
import time

def download_with_retry(url, local_path, max_retries=3):
    retries = 0
    while retries < max_retries:
        try:
            with urllib.request.urlopen(url) as response, open(local_path, 'wb') as fd:
                shutil.copyfileobj(response, fd)
            return True
        except IncompleteRead:
            retries += 1
            time.sleep(2)  # 间隔2秒后重试
        except Exception as e:
            send_notice_mail("Error downloading the file:", e)
            return False
    send_notice_mail("Max retries reached, download failed:", "")
    return False

额外验证步骤:

下载完成后,用zipfile校验文件完整性:

import zipfile
try:
    with zipfile.ZipFile(localfiletitle) as zf:
        # 检查损坏文件,返回None则无问题
        if zf.testzip() is not None:
            send_notice_mail("Downloaded zip file is corrupted", "")
            return False
except zipfile.BadZipFile:
    send_notice_mail("Downloaded file is not a valid zip", "")
    return False

内容的提问来源于stack exchange,提问作者Fabio Marzocca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 09:57:20