如何用Kotlin将Zip内容提取至单一扁平化目录
问题
我用Kotlin写了解析提取Zip文件的代码,已经能处理嵌套Zip归档,但现在需要实现新需求:把所有文件(无论嵌套层级)直接提取到单一目录中,不需要事后移动文件。当前代码会保留嵌套的目录结构,生成子目录存放文件,请问怎么修改代码实现需求?
现有代码:
internal fun File.unzipServiceFile(toPath: String): List<File> { val retFiles = mutableListOf<File>() val stream =ZipInputStream(this.inputStream()) stream.unzipServiceFile( toPath, retFiles) return retFiles } private fun ZipInputStream.unzipServiceFile( toPath: String, files:MutableList<File>) { var entry = nextEntry println("Parsing entry $entry") while( entry != null){ println("entry: " + entry.name) //if there are nested zip files, we need to extract them if (entry.name.endsWith(".zip")) { //we need to go deeper ZipInputStream(this).unzipServiceFile( toPath, files) } else if (entry.name.endsWith(".json") && !entry.isDirectory && !entry.name.startsWith(".") && !entry.name.startsWith( "_" ) ) { val file = File("$toPath/${entry.name}") FileUtils.writeByteArrayToFile(file, readBytes()) files.add(file) println("Added file ${file.path}") } entry = nextEntry } println() }
补充信息:
- 压缩前的文件夹结构:根目录下有多个子目录,每个子目录里有json文件,还有嵌套的Zip文件(子目录内的Zip里也包含json文件)。
- 当前代码输出结果:会生成和压缩包内一致的子目录结构,嵌套Zip里的json文件会被放到对应的子目录中。
解决方案
要实现所有文件平级提取到目标目录,需要做三个核心修改:
1. 丢弃目录路径,只保留文件名
不再使用entry.name作为完整文件路径,而是提取其中的纯文件名(即路径的最后一段)。比如subdir/file.json仅保留file.json。
2. 处理文件名冲突
如果不同层级的文件重名,直接覆盖会导致数据丢失,因此需要检测目标目录中是否已存在同名文件,若存在则添加后缀(如file_1.json、file_2.json)避免冲突。
3. 修复嵌套Zip的读取逻辑
原代码中ZipInputStream(this)的写法错误,会导致读取嵌套Zip时内容混乱。正确做法是先把当前ZipEntry的内容读入字节数组,再用ByteArrayInputStream创建新的ZipInputStream解析嵌套内容。
修改后的完整代码
import org.apache.commons.io.FileUtils import java.io.ByteArrayInputStream import java.io.File import java.util.zip.ZipInputStream internal fun File.unzipServiceFile(toPath: String): List<File> { val retFiles = mutableListOf<File>() ZipInputStream(this.inputStream()).use { stream -> stream.unzipServiceFile(toPath, retFiles) } return retFiles } private fun ZipInputStream.unzipServiceFile(toPath: String, files: MutableList<File>) { var entry = nextEntry while (entry != null) { when { // 处理嵌套Zip文件 entry.name.endsWith(".zip") && !entry.isDirectory -> { // 读取当前ZipEntry的字节内容,创建新流解析嵌套Zip val zipBytes = readBytes() ByteArrayInputStream(zipBytes).use { byteStream -> ZipInputStream(byteStream).use { nestedStream -> nestedStream.unzipServiceFile(toPath, files) } } } // 处理目标json文件(过滤目录、隐藏文件、下划线开头的文件) entry.name.endsWith(".json") && !entry.isDirectory && !entry.name.startsWith(".") && !entry.name.startsWith("_") -> { // 提取纯文件名,丢弃原始路径 val fileName = File(entry.name).name // 生成不冲突的目标文件 val targetFile = getUniqueFile(File(toPath), fileName) // 写入文件 FileUtils.writeByteArrayToFile(targetFile, readBytes()) files.add(targetFile) println("Added file ${targetFile.path}") } } entry = nextEntry } } /** * 生成目标目录中不重名的文件 */ private fun getUniqueFile(targetDir: File, fileName: String): File { val baseName = fileName.substringBeforeLast(".") val extension = fileName.substringAfterLast(".", "") var counter = 1 var targetFile = File(targetDir, fileName) while (targetFile.exists()) { // 生成带后缀的文件名,如 file_1.json val newFileName = if (extension.isEmpty()) { "$baseName_$counter" } else { "$baseName_$counter.$extension" } targetFile = File(targetDir, newFileName) counter++ } return targetFile }
关键修改说明
- 使用
File(entry.name).name提取纯文件名,彻底丢弃原始目录结构; - 新增
getUniqueFile函数处理重名文件,避免数据覆盖; - 嵌套Zip读取时,先读取当前entry的字节内容再创建新流,确保嵌套Zip内容被正确解析;
- 添加
use函数自动关闭流,避免资源泄漏。
内容的提问来源于stack exchange,提问作者pbuchheit
相关产品推荐
相关产品推荐

