You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从嵌套Zip文件中提取.mdf文件的Python代码问题

解决嵌套Zip中提取.mdf文件的路径错误问题

我有多个顶层Zip文件,内部包含嵌套Zip及其他类型文件,需要遍历顶层Zip,识别其中的嵌套Zip并提取内部的.mdf文件。原代码仅能提取嵌套Zip,修改后代码出现“No such file or directory: 'test.zip'”错误,相关代码如下:

原代码:

for file in os.listdir(working_directory):
    if zipfile.is_zipfile(file):
        with zipfile.ZipFile(file) as item:
            for member in item.namelist():  # go through members of the zip file
                if member.endswith('.zip'):
                    item.extract(member)    # extract only the mdf file

修改后报错代码:

for file in os.listdir(working_directory):
    if zipfile.is_zipfile(file):
        with zipfile.ZipFile(file) as item: 
            for member in item.namelist():  # go through members of the zip file
                if member.endswith('.zip'):
                    with zipfile.ZipFile(member) as item2: 
                        for member2 in item2.namelist(): 
                            if member2.endswith('.mdf'):
                                item2.extract(member2)

参考代码:

with zipfile.ZipFile('all.zip') as z:
    with z.open('nested.zip') as z2:
        z2_filedata =  io.BytesIO(z2.read())
        with zipfile.ZipFile(z2_filedata) as nested_zip:
            print( nested_zip.open('readme.md').read())

错误原因

修改后的代码直接用member作为路径创建ZipFile是错误的:member只是顶层Zip内部的嵌套Zip文件名,并没有被实际提取到本地(或者即使原代码有提取,也可能因为工作目录问题找不到)。参考代码的思路才是正确的——无需先将嵌套Zip提取到本地,直接在内存中读取其内容,既高效又不会产生临时文件。

修正后的完整代码

import os
import zipfile
import io

working_directory = "./your_working_dir"  # 替换为你的实际工作目录

for file in os.listdir(working_directory):
    # 拼接完整文件路径,避免相对路径导致的找不到文件问题
    full_top_zip_path = os.path.join(working_directory, file)
    if zipfile.is_zipfile(full_top_zip_path):
        with zipfile.ZipFile(full_top_zip_path) as top_zip:
            for nested_zip_name in top_zip.namelist():
                if nested_zip_name.endswith('.zip'):
                    # 直接读取嵌套Zip的字节流,在内存中处理
                    with top_zip.open(nested_zip_name) as nested_zip_stream:
                        nested_zip_data = io.BytesIO(nested_zip_stream.read())
                        with zipfile.ZipFile(nested_zip_data) as nested_zip:
                            # 遍历嵌套Zip,提取所有.mdf文件
                            for mdf_file in nested_zip.namelist():
                                if mdf_file.endswith('.mdf'):
                                    nested_zip.extract(mdf_file)

关键修改说明

  • 用os.path.join拼接顶层Zip的完整路径,解决os.listdir返回文件名无路径导致的找不到文件问题
  • 跳过提取嵌套Zip到本地的步骤,直接通过top_zip.open读取嵌套Zip的字节流,用io.BytesIO包装成可被ZipFile识别的类文件对象
  • 减少磁盘IO操作,避免临时文件堆积,同时彻底解决路径错误问题

内容的提问来源于stack exchange,提问作者CGarden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 21:30:59