Python从URL下载Zip文件提取失败,BadZipFile错误求助
问题描述
我有一个可下载Zip文件的URL:https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20(202405).zip,但编写的两段代码均无法正常工作,执行后都返回错误:
BadZipFile: File is not a zip file
手动打开下载后的文件也无法提取,怀疑下载过程中文件损坏,求解决办法。
两段出错代码如下:
第一段代码
from datetime import datetime, timedelta import requests, zipfile, io from zipfile import BadZipFile TwoMonthsAgo = datetime.now() - timedelta(60) zip_file_url = 'https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20('+ datetime.strftime(TwoMonthsAgo, '%Y%m') +').zip' r = requests.get(zip_file_url) z = zipfile.ZipFile(io.BytesIO(r.content)) z.extractall()
第二段代码
import urllib.request filename = 'P6420 Food and beverages ('+ datetime.strftime(TwoMonthsAgo, '%Y%m') +').zip' urllib.request.urlretrieve(zip_file_url, filename) import zipfile with zipfile.ZipFile(filename, 'r') as zip_ref: zip_ref.extractall()
解决思路及修正代码
核心问题分析
- 请求头缺失:很多服务器会校验
User-Agent,拒绝非浏览器发起的请求,返回的不是真实Zip文件(而是HTML错误页)。 - 日期计算不准确:
timedelta(60)简单减60天可能导致年月计算偏差,生成的URL对应文件不存在。 - 未校验响应状态:没检查请求是否成功,直接把错误页面当作Zip文件处理。
修正后的Requests版本代码
from datetime import datetime, timedelta import requests, zipfile, io from zipfile import BadZipFile # 更准确地计算两个月前的年月(避免跨月偏差) current = datetime.now() # 先取上个月最后一天,再取上上个月的当月 last_month_end = current.replace(day=1) - timedelta(days=1) two_months_ago = last_month_end.replace(day=1) - timedelta(days=1) date_str = two_months_ago.strftime('%Y%m') zip_file_url = f'https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20({date_str}).zip' # 添加浏览器模拟请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36' } # 发送请求并允许重定向 response = requests.get(zip_file_url, headers=headers, allow_redirects=True) # 先校验请求状态 if response.status_code != 200: print(f"请求失败,状态码:{response.status_code}") print("响应内容预览:", response.text[:500]) else: try: with zipfile.ZipFile(io.BytesIO(response.content)) as zip_ref: zip_ref.extractall() print("Zip文件提取成功") except BadZipFile: print("下载的内容不是有效Zip文件,可能URL错误或服务器返回非Zip内容") # 保存响应内容以便排查 with open('debug_response.html', 'wb') as f: f.write(response.content) print("已保存响应内容到debug_response.html,可打开查看具体错误")
修正后的Urllib版本代码
from datetime import datetime, timedelta import urllib.request import zipfile from zipfile import BadZipFile # 准确计算两个月前的年月 current = datetime.now() last_month_end = current.replace(day=1) - timedelta(days=1) two_months_ago = last_month_end.replace(day=1) - timedelta(days=1) date_str = two_months_ago.strftime('%Y%m') zip_file_url = f'https://www.statssa.gov.za/timeseriesdata/Excel/P6420%20Food%20and%20beverages%20({date_str}).zip' filename = f'P6420 Food and beverages ({date_str}).zip' # 添加请求头模拟浏览器 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36' } request = urllib.request.Request(zip_file_url, headers=headers) try: with urllib.request.urlopen(request) as response: with open(filename, 'wb') as f: f.write(response.read()) if response.getcode() != 200: print(f"请求失败,状态码:{response.getcode()}") else: try: with zipfile.ZipFile(filename, 'r') as zip_ref: zip_ref.extractall() print("Zip文件提取成功") except BadZipFile: print("下载的文件不是有效Zip,建议手动访问生成的URL确认文件是否存在") except Exception as e: print(f"下载过程出错:{str(e)}")
额外排查步骤
- 验证URL有效性:打印生成的
zip_file_url,手动在浏览器中访问,确认是否能正常下载Zip文件。如果手动也无法下载,说明该日期的文件不存在,需要调整日期逻辑。 - 查看响应内容:如果仍然报错,打开保存的
debug_response.html,查看服务器返回的具体错误信息(比如404、权限限制等)。
内容的提问来源于stack exchange,提问作者prashanth manohar
相关产品推荐
相关产品推荐

