使用RxJava+ZipOutputStream生成Zip并S3分段上传时内容损坏排查
问题:Zip文件分段上传至AWS S3后内容损坏,解压失败
我尝试创建Zip文件并通过AWS S3分段上传,但文件生成后内部内容损坏,解压时出现以下错误:
Archive: test2MB.zip End-of-central-directory signature not found. Either this file is not a zipfile, or it constitutes one disk of a multi-part archive. In the latter case the central directory and zipfile comment will be found on the last disk(s) of this archive. unzip: cannot find zipfile directory in one of test2MB.zip or test2MB.zip.zip, and cannot find test2MB.zip.ZIP, period.
相关代码片段如下:
主流程方法
private Single<List<CompletedPart>> zipAndUploadZipParts(String uploadId) { List<Single<CompletedPart>> uploadPartObservables = new ArrayList<>(); final int[] partNumber = {Constants.START_PART_NUMBER}; ByteArrayOutputStream byteArrayOutputStream = new ByteArrayOutputStream(); ZipOutputStream zipOutputStream = new ZipOutputStream(byteArrayOutputStream); return zipMetadata(zipOutputStream) .andThen(uploadZipPart(byteArrayOutputStream, uploadPartObservables, partNumber, uploadId) .onErrorResumeNext(error -> { this.contextRx.log().error(String.format("placeholder", partNumber[0], this.logIdentifier, error.getMessage()), error); return Completable.error(new RuntimeException("Placeholder")); })) .andThen(finalizeZipUpload(uploadPartObservables, zipOutputStream, byteArrayOutputStream)); }
压缩元数据方法
private Completable zipMetadata(ZipOutputStream zipOutputStream) { return Completable.create(emitter -> { try { ZipEntry metadataEntry = new ZipEntry(Constants.METADATA_FILE_NAME); zipOutputStream.putNextEntry(metadataEntry); zipOutputStream.write(this.catalogObjectsMetadata.getBytes()); zipOutputStream.closeEntry(); emitter.onComplete(); } catch (IOException exception) { this.contextRx.log().error(String.format(Constants.LOG_ERROR_ZIPPING_METADATA, this.logIdentifier, exception.getMessage()), exception); emitter.onError(new RuntimeException(String.format(Constants.ERROR_ZIPPING_METADATA))); } }); }
上传分段方法
private Completable uploadZipPart(ByteArrayOutputStream byteArrayOutputStream, List<Single<CompletedPart>> uploadPartObservables, int[] partNumber, String uploadId) { return Completable.create(emitter -> { try { byte[] zipData = byteArrayOutputStream.toByteArray(); uploadPartObservables.add(uploadPart(zipData, partNumber[0]++, uploadId)); byteArrayOutputStream.reset(); this.contextRx.log().info(String.format("Uploaded part %d for uploadId %s", partNumber[0] - 1, uploadId)); emitter.onComplete(); } catch (Exception exception) { emitter.onError(exception); } }); }
收尾上传方法
private Single<List<CompletedPart>> finalizeZipUpload(List<Single<CompletedPart>> uploadPartObservables, ZipOutputStream zipOutputStream, ByteArrayOutputStream byteArrayOutputStream) { return Single.defer(() -> { try { zipOutputStream.finish(); } catch (IOException e) { return Single.error(new RuntimeException("Error finishing zip output stream", e)); } finally { try { zipOutputStream.close(); } catch (IOException e) { return Single.error(new RuntimeException("Error closing zip output stream", e)); } try { byteArrayOutputStream.close(); } catch (IOException e) { return Single.error(new RuntimeException("Error closing byte array output stream", e)); } } return Single.merge(uploadPartObservables) .toList() .doOnSuccess(parts -> this.contextRx.log().info(String.format(Constants.ALL_PARTS_UPLOADED, this.logIdentifier, parts))) .onErrorReturnItem(new ArrayList<>()); }); }
问题原因分析
核心问题是Zip文件的核心目录(Central Directory)和结束签名未被上传,代码逻辑顺序错误:
- 先上传了仅包含元数据条目的Zip片段,之后才调用
zipOutputStream.finish()生成Zip的核心目录和结束签名,但这部分关键内容没有被作为最后一个分段上传到S3。 finalizeZipUpload方法仅合并了之前添加的分段上传任务,完全忽略了finish()后写入到byteArrayOutputStream的核心目录数据,导致S3上的Zip文件缺失关键结构,解压时无法识别。
修复方案
修改finalizeZipUpload方法,在完成Zip流收尾后,将剩余的核心目录数据作为最后一个分段上传,再合并所有分段:
修改后的收尾上传方法
private Single<List<CompletedPart>> finalizeZipUpload(List<Single<CompletedPart>> uploadPartObservables, ZipOutputStream zipOutputStream, ByteArrayOutputStream byteArrayOutputStream, String uploadId) { return Single.defer(() -> { try { zipOutputStream.finish(); // 上传最后一部分:Zip的核心目录和结束签名 byte[] finalZipData = byteArrayOutputStream.toByteArray(); if (finalZipData.length > 0) { // 用现有分段数量+1作为最后一个分段的编号 uploadPartObservables.add(uploadPart(finalZipData, uploadPartObservables.size() + 1, uploadId)); } } catch (IOException e) { return Single.error(new RuntimeException("Error finishing zip output stream", e)); } finally { try { zipOutputStream.close(); } catch (IOException e) { return Single.error(new RuntimeException("Error closing zip output stream", e)); } try { byteArrayOutputStream.close(); } catch (IOException e) { return Single.error(new RuntimeException("Error closing byte array output stream", e)); } } return Single.merge(uploadPartObservables) .toList() .doOnSuccess(parts -> this.contextRx.log().info(String.format(Constants.ALL_PARTS_UPLOADED, this.logIdentifier, parts))) .onErrorReturnItem(new ArrayList<>()); }); }
同步修改主流程的方法调用
在zipAndUploadZipParts中调用finalizeZipUpload时传入uploadId参数:
.andThen(finalizeZipUpload(uploadPartObservables, zipOutputStream, byteArrayOutputStream, uploadId));
内容的提问来源于stack exchange,提问作者Satyam Kumar
相关产品推荐
相关产品推荐

