使用Python请求库下载Zip文件仅获部分无效文件的问题求助
解决Zip文件下载不完整/IncompleteRead错误的方案
核心问题排查与解决步骤:
- 开启响应内容自动解码
requests的r.raw默认不会解码分块传输编码的内容,导致下载文件混入分块标记,直接损坏Zip结构。修改代码添加解码配置:
try: r = requests.get(BASEURL, stream=True) r.raw.decode_content = True # 关键:让raw自动处理分块编码 with open(localfiletitle, 'wb') as fd: shutil.copyfileobj(r.raw, fd) except requests.exceptions.RequestException as e: send_notice_mail("Error downloading the file:", e) return False
- 验证请求状态码
先确认服务器返回成功响应,避免下载错误页面而非目标文件:
try: r = requests.get(BASEURL, stream=True) r.raise_for_status() # 主动抛出4xx/5xx状态码的异常 r.raw.decode_content = True with open(localfiletitle, 'wb') as fd: shutil.copyfileobj(r.raw, fd) except requests.exceptions.RequestException as e: send_notice_mail("Error downloading the file:", e) return False
- 改用
iter_content分块读取
替代直接操作r.raw,使用requests提供的更稳定的分块读取方法:
try: r = requests.get(BASEURL, stream=True) r.raise_for_status() with open(localfiletitle, 'wb') as fd: # 按8KB块读取,可根据网络情况调整大小 for chunk in r.iter_content(chunk_size=8192): if chunk: # 过滤空块 fd.write(chunk) except requests.exceptions.RequestException as e: send_notice_mail("Error downloading the file:", e) return False
- 添加浏览器用户代理
部分服务器会拦截默认requests代理,返回不完整内容,模拟浏览器请求:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } try: r = requests.get(BASEURL, stream=True, headers=headers) r.raise_for_status() r.raw.decode_content = True with open(localfiletitle, 'wb') as fd: shutil.copyfileobj(r.raw, fd) except requests.exceptions.RequestException as e: send_notice_mail("Error downloading the file:", e) return False
- urllib的IncompleteRead异常处理
如果要继续使用urllib,可捕获异常并重试:
import urllib.request from http.client import IncompleteRead import time def download_with_retry(url, local_path, max_retries=3): retries = 0 while retries < max_retries: try: with urllib.request.urlopen(url) as response, open(local_path, 'wb') as fd: shutil.copyfileobj(response, fd) return True except IncompleteRead: retries += 1 time.sleep(2) # 间隔2秒后重试 except Exception as e: send_notice_mail("Error downloading the file:", e) return False send_notice_mail("Max retries reached, download failed:", "") return False
额外验证步骤:
下载完成后,用zipfile校验文件完整性:
import zipfile try: with zipfile.ZipFile(localfiletitle) as zf: # 检查损坏文件,返回None则无问题 if zf.testzip() is not None: send_notice_mail("Downloaded zip file is corrupted", "") return False except zipfile.BadZipFile: send_notice_mail("Downloaded file is not a valid zip", "") return False
内容的提问来源于stack exchange,提问作者Fabio Marzocca
相关产品推荐
相关产品推荐

