You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

嵌套Zip归档写入异常:Python检测到文件但7z不可见

问题

我有一个约200GB的大型Zip归档文件,其中包含多个子归档文件。操作根层级归档时,复制、移动文件等功能正常,但操作嵌套子归档时出现异常:可读取数据,但写入操作无报错却未生成实际文件。

归档结构示例:

Archive.zip
├── Folder1
│   ├── **Archive1_1.zip**
│   │   └── Folder1_1_1
│   │       └── stuff I have to work with...
│   ├── Archive1_2.zip
│   │   └── Folder1_2_1
│   │       └── stuff I have to work with...
│   └── Archive1_3.zip
│       └── another Folder1_3_1
│           └── stuff I have to work with...
├── Folder2
│   ├── Archive2_1.zip
│   │   └── Folder2_1_1
│   │       └── stuff I have to work with...
│   └── Folder2_2
│       └── stuff I have to work with...
└── Folder3
    └── Folder3_1
        └── stuff I have to work with...

比如向Archive1_1.zip内的./foo/bar/file_C.txt写入文件,原归档含2个文件、2个目录共4个条目,写入后zipfile显示条目数变为5,再次运行脚本也能检测到该文件,但用7z查看时却找不到。

相关代码:

归档打开方式

with zipfile.ZipFile(zipPath, mode='a') as root_archive:
    for file_name in root_archive.namelist():
        if re.search(r'\.zip$', file_name) is not None:
            zip_archive = BytesIO(root_archive.read(file_name))
            with zipfile.ZipFile(zip_archive, mode='a') as sub_archive:
                start(sub_archive)

    start(root_archive)

文件写入代码

config.zip_archive.write(f"./tmp/{ident}{constants.SUFFIX}", 
                         f"{path}/{ident}{constants.SUFFIX}", 
                         compress_type=ZIP_DEFLATED)

验证代码

print(f"Files before/after: {len(archive.filelist)}")

解决方案

问题根源

你犯了一个典型的内存操作 vs 持久化存储错误:你把嵌套的子归档读到了内存中的BytesIO对象里,修改后只在内存里的sub_archive对象中生效,但从来没把修改后的子归档内容写回根Zip文件。脚本运行时,内存里的对象能看到新增条目,但根归档里的原Archive1_1.zip完全没被更新,所以7z这类外部工具自然看不到变化。

修正步骤

要解决这个问题,你需要在内存中修改完子归档后,将更新后的内容替换回根归档里的原文件。注意:Python的zipfile模块在追加模式('a')下无法删除现有条目,所以需要通过复制+替换的方式重新构建根归档内容,具体代码如下:

import zipfile
import io
import re
import shutil
import tempfile
import os

zipPath = "path/to/your/Archive.zip"

# 处理200GB大文件,用临时文件而非纯内存操作,避免内存溢出
with tempfile.NamedTemporaryFile(delete=False) as temp_root_file:
    temp_root_path = temp_root_file.name

try:
    with zipfile.ZipFile(zipPath, 'r') as root_archive, \
         zipfile.ZipFile(temp_root_path, 'w', zipfile.ZIP_DEFLATED) as temp_archive:
        
        # 先复制根归档中所有不需要修改的条目
        for item in root_archive.infolist():
            if not item.filename.endswith('.zip'):
                temp_archive.writestr(item, root_archive.read(item.filename))
        
        # 处理每个子归档
        for file_name in root_archive.namelist():
            if not file_name.endswith('.zip'):
                continue
            
            # 读取子归档到内存
            sub_zip_data = root_archive.read(file_name)
            zip_archive = io.BytesIO(sub_zip_data)
            
            # 修改子归档(执行你的写入逻辑)
            with zipfile.ZipFile(zip_archive, 'a') as sub_archive:
                start(sub_archive)
            
            # 将修改后的子归档写入临时根归档
            zip_archive.seek(0)
            temp_archive.writestr(file_name, zip_archive.getvalue())
    
    # 用临时文件替换原根归档(操作前务必备份原文件!)
    shutil.move(temp_root_path, zipPath)
finally:
    # 清理临时文件(如果替换失败)
    if os.path.exists(temp_root_path):
        os.unlink(temp_root_path)

关键注意事项

  1. 大文件内存优化:200GB的归档直接用内存处理会导致内存溢出,改用临时文件中转更安全。
  2. zipfile局限性:zipfile的追加模式不支持删除条目,必须通过复制+替换的方式更新子归档,这是这类操作的标准流程。
  3. 数据安全:操作大文件前务必备份原归档,避免操作失败导致数据丢失。
  4. 验证方式:修改完成后,重新打开根归档读取子归档,或者用7z直接查看,确认新增文件存在。

内容的提问来源于stack exchange,提问作者CSharper96

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 00:35:30