Java ZipOutputStream生成损坏文件:CSV内容部分重复
Spring Boot 3中ZIP打包CSV文件损坏问题排查与解决
问题现象
Spring Boot 3项目中,使用Jackson CSV Mapper生成的CSV文件保存到磁盘后内容正常,但通过ZipOutputStream打包为ZIP后,提取出的CSV文件出现损坏:内容部分重复、部分行缺失前引号、末尾存在多余引号。
相关代码
CSV生成代码
private String collectValuesToCSV(List<? extends Object> values, CsvMapper csvMapper, CsvSchema csvSchema) { String result; try (StringWriter strW = new StringWriter(); SequenceWriter seqW = csvMapper.writer(csvSchema).writeValues(strW)) { for (Object value : values) { seqW.write(value); } seqW.flush(); strW.flush(); result = strW.toString(); } catch (IOException e) { log.fatal(String.format("%s - Unable to generate CSV", LOG_PREFIX), e); throw e; } return result; } // 使用示例 CsvMapper csvMapper = CsvMapper.builder() .enable(CsvGenerator.Feature.ALWAYS_QUOTE_STRINGS) .enable(CsvGenerator.Feature.ALWAYS_QUOTE_EMPTY_STRINGS) .build(); CsvSchema csvSchema = csvMapper .schemaFor(MyDTO.class) .withColumnSeparator(';') .withHeader(); List<MyDTO> objects = generateObjects(); String myCsv = collectValuesToCSV(objects, csvMapper, csvSchema); byte[] contents = myCsv.getBytes(StandardCharsets.UTF_8); Path filePath = Path.of(config.baseFolder, "output_" + dateStr + ".csv"); Files.write(filePath, contents);
ZIP打包代码
private void zipCsvFiles(String dateStr, Path path1, Path path2, Path path3) throws IOException { String zipFileName = Path.of(config.zipFolder, "output_" + dateStr + ".zip").toString(); try (FileOutputStream fos = new FileOutputStream(zipFileName); ZipOutputStream zipOut = new ZipOutputStream(fos)) { addZipEntry(zipOut, path1); addZipEntry(zipOut, path2); addZipEntry(zipOut, path3); } } private void addZipEntry(ZipOutputStream zipOut, Path path) throws IOException { try (FileInputStream fis = new FileInputStream(path.toFile())) { ZipEntry entry = new ZipEntry(path.getFileName().toString()); zipOut.putNextEntry(entry); byte[] bytes = new byte[1024]; while ((fis.read(bytes)) >= 0) { zipOut.write(bytes); } zipOut.closeEntry(); } }
故障原因
问题出在ZIP打包的addZipEntry方法中:
- 每次调用
fis.read(bytes)时,返回值是实际读取的字节数(最多1024),但代码忽略了这个返回值,直接写入整个1024字节的数组。 - 当文件内容长度不是1024的整数倍时,最后一次读取只会填满数组的一部分,剩余部分是之前读取的残留数据。写入整个数组会把这些残留数据附加到ZIP内的文件末尾,导致内容损坏(重复、引号异常等)。
解决方案
修改addZipEntry方法,记录每次读取的实际字节数,并只写入对应长度的内容:
private void addZipEntry(ZipOutputStream zipOut, Path path) throws IOException { try (FileInputStream fis = new FileInputStream(path.toFile())) { ZipEntry entry = new ZipEntry(path.getFileName().toString()); zipOut.putNextEntry(entry); byte[] bytes = new byte[1024]; int readLen; while ((readLen = fis.read(bytes)) != -1) { zipOut.write(bytes, 0, readLen); } zipOut.closeEntry(); } }
额外优化(可选)
如果使用Java 9+,可以直接使用InputStream.transferTo简化代码,避免手动处理字节数组:
private void addZipEntry(ZipOutputStream zipOut, Path path) throws IOException { try (FileInputStream fis = new FileInputStream(path.toFile())) { ZipEntry entry = new ZipEntry(path.getFileName().toString()); zipOut.putNextEntry(entry); fis.transferTo(zipOut); zipOut.closeEntry(); } }
内容的提问来源于stack exchange,提问作者Gábor Major
相关产品推荐
相关产品推荐

