Python 3下无法生成有效Zip文件:解压时出现中央目录签名缺失错误
看起来你遇到的麻烦是:相同的代码在Python2里能生成正常解压的Zip文件,但到了Python3环境,生成的storage_aligned.zip就成了无效文件,解压时会抛出这个错误:
Archive: storage_aligned.zip
End-of-central-directory signature not found. Either this file is not a zipfile, or it constitutes one disk of a multi-part archive. In the latter case the central directory and zipfile comment will be found on the last disk(s) of this archive.
unzip: cannot find zipfile directory in one of storage_aligned.zip or storage_aligned.zip.zip, and cannot find storage_aligned.zip.ZIP, period
问题根源
核心问题出在你直接操作了ZipFile的底层文件指针(aligned_zf.fp)。Python2的文件对象缓冲逻辑比较宽松,直接写入额外字节刚好没破坏Zip结构;但Python3的zipfile模块对文件指针的管理更严谨,你绕开API直接写填充字节的行为,会打乱模块内部维护的文件偏移量——当调用aligned_zf.close()时,中央目录和结尾签名会被写到错误的位置,最终导致Zip文件结构损坏。
修复方案
不要直接操作ZipFile的fp属性,改用手动构建Zip内容+复用合法中央目录的方式,既保留你的填充逻辑,又保证Zip结构的有效性。修复后的代码如下:
import zipfile import struct import sys import io # 假设self是原ZipFile实例,file是输出文件路径 padding_index = 0 # 假设该变量已初始化 paddings = [] # 假设该列表已填充合法的填充长度值 # 第一步:构建带填充的文件条目内容(LFH+文件名+内容+填充) output_content = b"" for zinfo in self.filelist: zef_file = self.fp zef_file.seek(zinfo.header_offset, 0) # 读取原文件头 fheader = zef_file.read(zipfile.sizeFileHeader) if fheader[:4] != zipfile.stringFileHeader: raise zipfile.BadZipfile("Bad magic number for file header") fheader_unpacked = struct.unpack(zipfile.structFileHeader, fheader) fname_bytes = zef_file.read(fheader_unpacked[zipfile._FH_FILENAME_LENGTH]) # 读取额外字段(如果有) extra_bytes = b"" if fheader_unpacked[zipfile._FH_EXTRA_FIELD_LENGTH]: extra_bytes = zef_file.read(fheader_unpacked[zipfile._FH_EXTRA_FIELD_LENGTH]) # 读取原文件内容 file_content = self.read(fname_bytes.decode('utf-8')) # 写入前置填充(仅第一个条目) if padding_index == 0: output_content += b"\0" * paddings[padding_index] padding_index += 1 # 写入文件头、文件名、额外字段、文件内容 output_content += fheader output_content += fname_bytes output_content += extra_bytes output_content += file_content # 写入后置填充 output_content += b"\0" * paddings[padding_index] padding_index += 1 # 第二步:生成合法的中央目录和结尾签名 temp_zip = io.BytesIO() temp_zf = zipfile.ZipFile(temp_zip, 'w') for zinfo in self.filelist: temp_zf.writestr(zinfo, self.read(zinfo.orig_filename)) temp_zf.close() # 提取中央目录部分(从临时Zip中获取合法的结构) temp_zip.seek(0) temp_content = temp_zip.read() eocd_signature = b"\x50\x4b\x05\x06" # Zip结尾签名 eocd_start = temp_content.rfind(eocd_signature) central_dir = temp_content[eocd_start:] # 第三步:合并内容并写入最终文件 final_content = output_content + central_dir with open(file, 'wb') as f: f.write(final_content)
关键修复细节
- 放弃直接操作
fp:Python3的zipfile模块不允许绕开API修改底层文件,否则会破坏内部偏移量管理。 - 手动构建条目内容:把需要的填充字节和原文件的LFH、内容合并,确保这部分结构符合Zip规范。
- 复用合法中央目录:通过临时Zip生成标准的中央目录和结尾签名,再拼接到我们的内容后面——这保证了Zip文件的核心结构是有效的,解压工具能正确识别。
- 统一字节处理:全程使用二进制字节串(
b"")操作,避免Python3中字符串和字节串的编码混淆问题。
这样修改后,Python3环境下生成的Zip文件就能正常解压了,同时也兼容Python2。
内容的提问来源于stack exchange,提问作者gilly

