You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Kotlin提取包含嵌套Zip文件的压缩包

解决方案

你可以通过抽离统一的Zip流处理逻辑实现递归提取嵌套Zip中的JSON文件,无需将中间Zip文件写入磁盘。核心思路是直接在内存中处理嵌套Zip的输入流,递归解析其中的内容:

修改后的完整代码

import org.apache.commons.io.FileUtils
import java.io.ByteArrayInputStream
import java.io.File
import java.io.InputStream
import java.util.zip.ZipEntry
import java.util.zip.ZipInputStream

fun File.unzipServiceFile(toPath: String): List<File> {
    val extractedFiles = mutableListOf<File>()
    this.inputStream().use { fileStream ->
        extractedFiles.addAll(processZipStream(fileStream, toPath))
    }
    return extractedFiles
}

private fun processZipStream(inputStream: InputStream, toPath: String): List<File> {
    val resultFiles = mutableListOf<File>()
    ZipInputStream(inputStream).use { zipStream ->
        var currentEntry: ZipEntry? = zipStream.nextEntry
        while (currentEntry != null) {
            if (!currentEntry.isDirectory) {
                when {
                    // 处理嵌套Zip文件:读取流内容递归解析
                    currentEntry.name.endsWith(".zip") -> {
                        val zipBytes = zipStream.readBytes()
                        ByteArrayInputStream(zipBytes).use { nestedZipStream ->
                            resultFiles.addAll(processZipStream(nestedZipStream, toPath))
                        }
                    }
                    // 处理符合条件的JSON文件
                    currentEntry.name.endsWith(".json") 
                    && !currentEntry.name.startsWith(".") 
                    && !currentEntry.name.startsWith("_") -> {
                        val targetFile = File("$toPath/${currentEntry.name}")
                        // 确保目标文件的父目录存在
                        targetFile.parentFile?.mkdirs()
                        // 流复制写入文件,比readBytes()更节省内存
                        targetFile.outputStream().use { output ->
                            zipStream.copyTo(output)
                        }
                        resultFiles.add(targetFile)
                    }
                }
            }
            zipStream.closeEntry()
            currentEntry = zipStream.nextEntry
        }
    }
    return resultFiles
}

关键实现要点

  • 统一流处理:将Zip解析逻辑抽离为processZipStream函数,既可以处理磁盘上的Zip文件流,也可以处理嵌套Zip条目流,实现递归复用。
  • 内存中处理嵌套Zip:遇到嵌套Zip条目时,直接读取其字节内容到内存,转换为ByteArrayInputStream后递归解析,完全避免中间Zip文件落地。
  • 资源自动管理:所有流都通过use块自动关闭,避免资源泄漏。
  • 目录创建:写入JSON文件前先创建父目录,避免因路径不存在导致的写入失败。
  • 高效流复制:使用copyTo方法替代readBytes(),减少内存占用,适配更大的文件场景。

内容的提问来源于stack exchange,提问作者pbuchheit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 21:54:57