如何用Python从ZIP文件提取CSV?下载解压遇BadZipFile错误
解决下载并解压Zip文件时的BadZipFile错误
问题场景
你作为Python新手,尝试用以下代码从URL下载.zip文件并提取CSV,但运行后报错:
原代码
# importing necessary modules import requests, zipfile from io import BytesIO print('Downloading started') #Defining the zip file URL url = 'https://enfxfr.dol.gov/data_catalog/MSHA/msha_accident_20230422.csv.zip' # Split URL to get the file name filename = url.split('/')[-1] # Downloading the file by sending the request to the URL req = requests.get(url) print('Downloading Completed') # extracting the zip file contents zipfile= zipfile.ZipFile(BytesIO(req.content)) zipfile.extractall('C:/Users/ssawant/OneDrive - Iron Senergy/Desktop/MSHA converted data/msha_accident')
错误信息
Traceback (most recent call last): File "c:\Users\ssawant\OneDrive - Iron Senergy\Desktop\MSHA converted data\test_2.py", line 17, in <module> zipfile= zipfile.ZipFile(BytesIO(req.content)) File "C:\Users\ssawant\AppData\Local\Programs\Python\Python39\lib\zipfile.py", line 1266, in __init__ self._RealGetContents() File "C:\Users\ssawant\AppData\Local\Programs\Python\Python39\lib\zipfile.py", line 1333, in _RealGetContents raise BadZipFile("File is not a zip file") zipfile.BadZipFile: File is not a zip file
问题原因
- 变量名覆盖模块名:用
zipfile作为变量名,直接覆盖了导入的zipfile模块,会干扰模块方法的调用,导致ZipFile对象初始化异常。 - 未验证请求有效性:直接使用
req.content,如果请求返回错误页面(如404、500状态),得到的内容并非有效Zip文件,触发BadZipFile错误。
修复方案
1. 修改冲突变量名
将存储ZipFile对象的变量名改为zip_ref或其他不与模块名重复的名称。
2. 添加请求状态校验
使用req.raise_for_status()检查请求是否成功,失败时直接抛出异常,避免处理无效内容。
3. 可选:先保存本地文件
若仍有问题,可先将下载内容保存为本地文件,方便排查是否为有效Zip。
修复后的完整代码
import requests import zipfile from io import BytesIO print('Downloading started') url = 'https://enfxfr.dol.gov/data_catalog/MSHA/msha_accident_20230422.csv.zip' filename = url.split('/')[-1] # 发送请求并校验状态 req = requests.get(url) req.raise_for_status() # 请求失败时直接抛出异常 print('Downloading Completed') # 避免变量名覆盖模块 zip_ref = zipfile.ZipFile(BytesIO(req.content)) extract_path = 'C:/Users/ssawant/OneDrive - Iron Senergy/Desktop/MSHA converted data/msha_accident' zip_ref.extractall(extract_path) zip_ref.close() # 关闭Zip文件对象 print(f"文件已成功解压到 {extract_path}")
额外排查建议
如果还是报错,先保存下载内容到本地排查:
# 请求成功后添加这段代码保存文件 with open(filename, 'wb') as f: f.write(req.content)
手动打开文件,若为HTML页面则说明URL失效或网站需要验证;若为损坏的Zip,尝试分块下载:
req = requests.get(url, stream=True) req.raise_for_status() with open(filename, 'wb') as f: for chunk in req.iter_content(chunk_size=8192): f.write(chunk) # 从本地文件解压 zip_ref = zipfile.ZipFile(filename) zip_ref.extractall(extract_path) zip_ref.close()
内容的提问来源于stack exchange,提问作者So50935
相关产品推荐
相关产品推荐

