You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从URL下载Zip文件提取失败,BadZipFile错误求助

问题描述

我有一个可下载Zip文件的URL:https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20(202405).zip,但编写的两段代码均无法正常工作,执行后都返回错误:

BadZipFile: File is not a zip file

手动打开下载后的文件也无法提取,怀疑下载过程中文件损坏,求解决办法。

两段出错代码如下:

第一段代码

from datetime import datetime, timedelta
import requests, zipfile, io
from zipfile import BadZipFile

TwoMonthsAgo = datetime.now() - timedelta(60)
zip_file_url = 'https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20('+ datetime.strftime(TwoMonthsAgo, '%Y%m') +').zip'

r = requests.get(zip_file_url)
z = zipfile.ZipFile(io.BytesIO(r.content))
z.extractall()

第二段代码

import urllib.request 
filename = 'P6420 Food and beverages ('+ datetime.strftime(TwoMonthsAgo, '%Y%m') +').zip'
urllib.request.urlretrieve(zip_file_url, filename)

import zipfile
with zipfile.ZipFile(filename, 'r') as zip_ref:
    zip_ref.extractall()

解决思路及修正代码

核心问题分析

  1. 请求头缺失:很多服务器会校验User-Agent,拒绝非浏览器发起的请求,返回的不是真实Zip文件(而是HTML错误页)。
  2. 日期计算不准确:timedelta(60)简单减60天可能导致年月计算偏差,生成的URL对应文件不存在。
  3. 未校验响应状态:没检查请求是否成功,直接把错误页面当作Zip文件处理。

修正后的Requests版本代码

from datetime import datetime, timedelta
import requests, zipfile, io
from zipfile import BadZipFile

# 更准确地计算两个月前的年月(避免跨月偏差)
current = datetime.now()
# 先取上个月最后一天,再取上上个月的当月
last_month_end = current.replace(day=1) - timedelta(days=1)
two_months_ago = last_month_end.replace(day=1) - timedelta(days=1)
date_str = two_months_ago.strftime('%Y%m')

zip_file_url = f'https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20({date_str}).zip'

# 添加浏览器模拟请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36'
}

# 发送请求并允许重定向
response = requests.get(zip_file_url, headers=headers, allow_redirects=True)

# 先校验请求状态
if response.status_code != 200:
    print(f"请求失败,状态码:{response.status_code}")
    print("响应内容预览:", response.text[:500])
else:
    try:
        with zipfile.ZipFile(io.BytesIO(response.content)) as zip_ref:
            zip_ref.extractall()
            print("Zip文件提取成功")
    except BadZipFile:
        print("下载的内容不是有效Zip文件,可能URL错误或服务器返回非Zip内容")
        # 保存响应内容以便排查
        with open('debug_response.html', 'wb') as f:
            f.write(response.content)
        print("已保存响应内容到debug_response.html,可打开查看具体错误")

修正后的Urllib版本代码

from datetime import datetime, timedelta
import urllib.request
import zipfile
from zipfile import BadZipFile

# 准确计算两个月前的年月
current = datetime.now()
last_month_end = current.replace(day=1) - timedelta(days=1)
two_months_ago = last_month_end.replace(day=1) - timedelta(days=1)
date_str = two_months_ago.strftime('%Y%m')

zip_file_url = f'https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20({date_str}).zip'
filename = f'P6420 Food and beverages ({date_str}).zip'

# 添加请求头模拟浏览器
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36'
}
request = urllib.request.Request(zip_file_url, headers=headers)

try:
    with urllib.request.urlopen(request) as response:
        with open(filename, 'wb') as f:
            f.write(response.read())
        
        if response.getcode() != 200:
            print(f"请求失败,状态码:{response.getcode()}")
        else:
            try:
                with zipfile.ZipFile(filename, 'r') as zip_ref:
                    zip_ref.extractall()
                    print("Zip文件提取成功")
            except BadZipFile:
                print("下载的文件不是有效Zip,建议手动访问生成的URL确认文件是否存在")
except Exception as e:
    print(f"下载过程出错:{str(e)}")

额外排查步骤

  1. 验证URL有效性:打印生成的zip_file_url,手动在浏览器中访问,确认是否能正常下载Zip文件。如果手动也无法下载,说明该日期的文件不存在,需要调整日期逻辑。
  2. 查看响应内容:如果仍然报错,打开保存的debug_response.html,查看服务器返回的具体错误信息(比如404、权限限制等)。

内容的提问来源于stack exchange,提问作者prashanth manohar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 07:49:59