嵌套Zip归档写入异常:Python检测到文件但7z不可见
问题
我有一个约200GB的大型Zip归档文件,其中包含多个子归档文件。操作根层级归档时,复制、移动文件等功能正常,但操作嵌套子归档时出现异常:可读取数据,但写入操作无报错却未生成实际文件。
归档结构示例:
Archive.zip ├── Folder1 │ ├── **Archive1_1.zip** │ │ └── Folder1_1_1 │ │ └── stuff I have to work with... │ ├── Archive1_2.zip │ │ └── Folder1_2_1 │ │ └── stuff I have to work with... │ └── Archive1_3.zip │ └── another Folder1_3_1 │ └── stuff I have to work with... ├── Folder2 │ ├── Archive2_1.zip │ │ └── Folder2_1_1 │ │ └── stuff I have to work with... │ └── Folder2_2 │ └── stuff I have to work with... └── Folder3 └── Folder3_1 └── stuff I have to work with...
比如向Archive1_1.zip内的./foo/bar/file_C.txt写入文件,原归档含2个文件、2个目录共4个条目,写入后zipfile显示条目数变为5,再次运行脚本也能检测到该文件,但用7z查看时却找不到。
相关代码:
归档打开方式
with zipfile.ZipFile(zipPath, mode='a') as root_archive: for file_name in root_archive.namelist(): if re.search(r'\.zip$', file_name) is not None: zip_archive = BytesIO(root_archive.read(file_name)) with zipfile.ZipFile(zip_archive, mode='a') as sub_archive: start(sub_archive) start(root_archive)
文件写入代码
config.zip_archive.write(f"./tmp/{ident}{constants.SUFFIX}", f"{path}/{ident}{constants.SUFFIX}", compress_type=ZIP_DEFLATED)
验证代码
print(f"Files before/after: {len(archive.filelist)}")
解决方案
问题根源
你犯了一个典型的内存操作 vs 持久化存储错误:你把嵌套的子归档读到了内存中的BytesIO对象里,修改后只在内存里的sub_archive对象中生效,但从来没把修改后的子归档内容写回根Zip文件。脚本运行时,内存里的对象能看到新增条目,但根归档里的原Archive1_1.zip完全没被更新,所以7z这类外部工具自然看不到变化。
修正步骤
要解决这个问题,你需要在内存中修改完子归档后,将更新后的内容替换回根归档里的原文件。注意:Python的zipfile模块在追加模式('a')下无法删除现有条目,所以需要通过复制+替换的方式重新构建根归档内容,具体代码如下:
import zipfile import io import re import shutil import tempfile import os zipPath = "path/to/your/Archive.zip" # 处理200GB大文件,用临时文件而非纯内存操作,避免内存溢出 with tempfile.NamedTemporaryFile(delete=False) as temp_root_file: temp_root_path = temp_root_file.name try: with zipfile.ZipFile(zipPath, 'r') as root_archive, \ zipfile.ZipFile(temp_root_path, 'w', zipfile.ZIP_DEFLATED) as temp_archive: # 先复制根归档中所有不需要修改的条目 for item in root_archive.infolist(): if not item.filename.endswith('.zip'): temp_archive.writestr(item, root_archive.read(item.filename)) # 处理每个子归档 for file_name in root_archive.namelist(): if not file_name.endswith('.zip'): continue # 读取子归档到内存 sub_zip_data = root_archive.read(file_name) zip_archive = io.BytesIO(sub_zip_data) # 修改子归档(执行你的写入逻辑) with zipfile.ZipFile(zip_archive, 'a') as sub_archive: start(sub_archive) # 将修改后的子归档写入临时根归档 zip_archive.seek(0) temp_archive.writestr(file_name, zip_archive.getvalue()) # 用临时文件替换原根归档(操作前务必备份原文件!) shutil.move(temp_root_path, zipPath) finally: # 清理临时文件(如果替换失败) if os.path.exists(temp_root_path): os.unlink(temp_root_path)
关键注意事项
- 大文件内存优化:200GB的归档直接用内存处理会导致内存溢出,改用临时文件中转更安全。
- zipfile局限性:
zipfile的追加模式不支持删除条目,必须通过复制+替换的方式更新子归档,这是这类操作的标准流程。 - 数据安全:操作大文件前务必备份原归档,避免操作失败导致数据丢失。
- 验证方式:修改完成后,重新打开根归档读取子归档,或者用7z直接查看,确认新增文件存在。
内容的提问来源于stack exchange,提问作者CSharper96
相关产品推荐
相关产品推荐

