You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

嵌套Zip文件中提取PDF失败问题求助

解决嵌套Zip文件中PDF提取的问题

问题背景

需要提取嵌套在Zip文件内的Zip文件中的PDF文件,但现有代码要么跳过嵌套内的PDF,要么抛出java.io.IOException: Stream Closed错误,全网及Stack Overflow未找到匹配解决方案。

原代码问题分析

  • 流冲突:直接复用外层ZipInputStream创建内层ZipInputStream,导致外层流被内层操作破坏,后续关闭entry时触发流关闭错误。
  • 判断逻辑错误:内层循环中判断的是外层ZipEntry的文件名是否为PDF,而非内层entry,导致永远不会触发PDF写入逻辑。
  • 路径处理混乱:未正确创建内层Zip的临时存储,直接复用外层文件路径,导致文件操作逻辑错误。

修复后的完整代码

import java.io.*;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.zip.ZipEntry;
import java.util.zip.ZipInputStream;

public class NestedZipExtractor {
    private static final String ZIP_EXTENSION = ".ZIP";
    private static final String PDF_EXTENSION = ".PDF";
    private static final int BUFFER_SIZE = 4096;
    private static final String DESTINATION_DIRECTORY = "C:\\Users\\user\\Desktop\\Scan\\Extracted\\";

    public static void main(String[] args) {
        try {
            String basePath = "C:\\Users\\user\\Desktop\\Scan\\";
            File lookupDir = new File(basePath + "Data\\");
            String doneFolder = basePath + "DoneUnzipping\\";

            // 创建目标目录
            new File(doneFolder).mkdirs();
            new File(DESTINATION_DIRECTORY).mkdirs();

            File[] directoryListing = lookupDir.listFiles();
            if (directoryListing == null) {
                System.out.println("指定目录为空或无法访问");
                return;
            }

            for (File file : directoryListing) {
                if (file.isFile() && file.getName().toUpperCase().endsWith(ZIP_EXTENSION)) {
                    String pathOrigFile = file.getAbsolutePath();
                    Path origFileDone = Paths.get(pathOrigFile);
                    Path newFileDone = Paths.get(doneFolder + file.getName());

                    // 递归解压所有嵌套Zip中的PDF
                    unzip(file.getAbsolutePath(), DESTINATION_DIRECTORY);

                    // 将原Zip移动到Done文件夹
                    Files.move(origFileDone, newFileDone);
                }
            }
        } catch (Exception e) {
            e.printStackTrace(System.out);
        }
    }

    private static void unzip(String zipFilePath, String destDir) throws IOException {
        byte[] buffer = new byte[BUFFER_SIZE];

        try (ZipInputStream zis = new ZipInputStream(new FileInputStream(zipFilePath))) {
            ZipEntry ze;
            while ((ze = zis.getNextEntry()) != null) {
                String entryName = ze.getName();
                File entryFile = new File(destDir + File.separator + entryName);

                // 确保父目录存在
                if (!entryFile.getParentFile().exists()) {
                    entryFile.getParentFile().mkdirs();
                }

                if (ze.isDirectory()) {
                    entryFile.mkdirs();
                    zis.closeEntry();
                    continue;
                }

                // 如果是Zip文件,先写入临时文件再递归解压
                if (entryName.toUpperCase().endsWith(ZIP_EXTENSION)) {
                    try (FileOutputStream fos = new FileOutputStream(entryFile)) {
                        int len;
                        while ((len = zis.read(buffer)) > 0) {
                            fos.write(buffer, 0, len);
                        }
                    }
                    // 递归解压内层Zip
                    unzip(entryFile.getAbsolutePath(), destDir);
                    // 可选:删除临时Zip文件
                    entryFile.delete();
                }
                // 如果是PDF文件,直接写入目标目录
                else if (entryName.toUpperCase().endsWith(PDF_EXTENSION)) {
                    try (FileOutputStream fos = new FileOutputStream(entryFile)) {
                        int len;
                        while ((len = zis.read(buffer)) > 0) {
                            fos.write(buffer, 0, len);
                        }
                    }
                }

                zis.closeEntry();
            }
        }
    }
}

修复关键点说明

  • 递归处理嵌套Zip:遇到Zip文件时,先将其写入临时文件,再递归调用unzip方法,支持任意层数的嵌套。
  • 避免流冲突:不再复用外层ZipInputStream,而是将内层Zip写入临时文件后独立处理,彻底解决流关闭错误。
  • 修正判断逻辑:直接判断当前处理的ZipEntry是否为PDF,确保嵌套内的PDF被正确识别和写入。
  • 完善目录创建:自动创建目标目录和父目录,避免文件路径不存在的错误。
  • 自动流管理:使用try-with-resources语法自动关闭所有流,避免手动关闭导致的资源泄漏。
  • 可选清理临时文件:解压内层Zip后可删除临时文件,保持目标目录整洁。

内容的提问来源于stack exchange,提问作者user9604262

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 20:05:37