使用Pandas合并Excel文件时触发BadZipFile报错:File is not a zip file
解决Pandas合并Excel文件时的BadZipFile错误
我尝试用Pandas批量读取指定目录下的Excel文件,指定读取工作表cycle的Chg. Cap.(mAh)和DChg. Cap.(mAh)两列,合并到DataFrame中,但运行时触发BadZipFile: File is not a zip file错误。
原代码
import pandas as pd import glob df = [] list_of_df = pd.DataFrame() for filename in all_files: df = pd.read_excel(filename, sheet_name = 'cycle', usecols=['Chg. Cap.(mAh)','DChg. Cap.(mAh)'], engine = 'openpyxl') list_of_df = pd.concat([list_of_df. df], ignore_index = True) print(list_of_df) with pd.ExcelWriter('output3.xlsx', engine = 'openpyxl', mode = 'a',) as writer: df.to_excel(writer, sheet_name = 'output3')
报错栈
Traceback (most recent call last): File ~\AppData\Local\anaconda3\Lib\site-packages\spyder_kernels\py3compat.py:356 in compat_exec exec(code, globals, locals) File c:\documents\python scripts\chg._and_dchg._curves_for_excel5.py:30 df = pd.read_excel(filename, sheet_name = 'cycle', usecols=['Chg. Cap.(mAh)','DChg. Cap.(mAh)'], engine = 'openpyxl') File ~\AppData\Local\anaconda3\Lib\site-packages\pandas\util\_decorators.py:211 in wrapper return func(*args, **kwargs) File ~\AppData\Local\anaconda3\Lib\site-packages\pandas\util\_decorators.py:331 in wrapper return func(*args, **kwargs) File ~\AppData\Local\anaconda3\Lib\site-packages\pandas\io\excel\_base.py:482 in read_excel io = ExcelFile(io, storage_options=storage_options, engine=engine) File ~\AppData\Local\anaconda3\Lib\site-packages\pandas\io\excel\_base.py:1695 in __init__ self._reader = self._engines[engine](self._io, storage_options=storage_options) File ~\AppData\Local\anaconda3\Lib\site-packages\pandas\io\excel\_openpyxl.py:557 in __init__ super().__init__(filepath_or_buffer, storage_options=storage_options) File ~\AppData\Local\anaconda3\Lib\site-packages\pandas\io\excel\_base.py:545 in __init__ self.book = self.load_workbook(self.handles.handle) File ~\AppData\Local\anaconda3\Lib\site-packages\pandas\io\excel\_openpyxl.py:568 in load_workbook return load_workbook( File ~\AppData\Local\anaconda3\Lib\site-packages\openpyxl\reader\excel.py:315 in load_workbook reader = ExcelReader(filename, read_only, keep_vba, File ~\AppData\Local\anaconda3\Lib\site-packages\openpyxl\reader\excel.py:124 in __init__ self.archive = _validate_archive(fn) File ~\AppData\Local\anaconda3\Lib\site-packages\openpyxl\reader\excel.py:96 in _validate_archive archive = ZipFile(filename, 'r') File ~\AppData\Local\anaconda3\Lib\zipfile.py:1302 in __init__ self._RealGetContents() File ~\AppData\Local\anaconda3\Lib\zipfile.py:1369 in _RealGetContents raise BadZipFile("File is not a zip file") BadZipFile: File is not a zip file
错误原因及修复步骤
1. 排查目标文件合法性
BadZipFile错误核心是openpyxl无法解析目标文件,常见原因:
all_files包含非Excel文件(比如临时缓存文件、txt、csv等)- Excel文件损坏、未正常保存,或是旧版
.xls格式(openpyxl仅支持.xlsx/.xlsm等基于zip的格式)
修复操作:
- 用
glob精准指定文件后缀,只读取合法Excel文件:all_files = glob.glob("你的目标目录/*.xlsx") + glob.glob("你的目标目录/*.xlsm") - 手动遍历
all_files列表,删除损坏或非Excel文件
2. 修复代码语法错误
原代码中pd.concat([list_of_df. df])是笔误,点号需改为逗号:
list_of_df = pd.concat([list_of_df, df], ignore_index=True)
3. 优化合并逻辑(可选)
循环中反复concat效率较低,建议先将所有DataFrame存入列表,最后一次性合并:
import pandas as pd import glob # 指定目标目录和合法文件类型 all_files = glob.glob("your_directory/*.xlsx") + glob.glob("your_directory/*.xlsm") df_list = [] for filename in all_files: try: # 捕获单个文件读取错误,避免整个循环中断 df = pd.read_excel(filename, sheet_name='cycle', usecols=['Chg. Cap.(mAh)','DChg. Cap.(mAh)'], engine='openpyxl') # 添加来源文件名列,方便后续数据溯源(可选) df['source_file'] = filename df_list.append(df) except Exception as e: print(f"读取文件 {filename} 失败: {str(e)}") # 一次性合并所有DataFrame list_of_df = pd.concat(df_list, ignore_index=True) # 写入输出文件 with pd.ExcelWriter('output3.xlsx', engine='openpyxl') as writer: list_of_df.to_excel(writer, sheet_name='output3', index=False)
4. 处理旧版.xls文件
如果目录中有.xls格式文件,需改用xlrd引擎(注意xlrd 2.0+不再支持.xls,需安装xlrd==1.2.0):
# 针对.xls文件的读取代码 df = pd.read_excel(filename, sheet_name='cycle', usecols=['Chg. Cap.(mAh)','DChg. Cap.(mAh)'], engine='xlrd')
内容的提问来源于stack exchange,提问作者Stavo
相关产品推荐
相关产品推荐

