S3本地文件校验和的校验和与远程SHA-256不匹配问题排查
背景
我有一个约2.5GB的大文件存储在AWS S3中,上传时使用SHA-256作为校验和函数:
之后我参考AWS官方用户指南《检查对象完整性》,从「使用AWS SDK」部分复制了validateExistingFileAgainstS3Checksum函数,并做了以下修改:
- 根据IDE建议将部分代码提取为新函数
getPartBreak; - 移除部分
System.out.print语句; - 添加自定义日志以便排查问题;
- 重构错误处理逻辑;
- 将
getPartBreak返回的long类型强制转换为int类型以通过编译,确认该值不会超出int范围。
代码
以下是我修改后的代码:
package app.service.vendor.amazon; import app.service.vendor.amazon.exception.ChecksumValidationException; import io.netty.handler.codec.base64.Base64Encoder; import jakarta.inject.Inject; import jakarta.inject.Singleton; import org.slf4j.Logger; import software.amazon.awssdk.services.s3.S3Client; import software.amazon.awssdk.services.s3.model.GetObjectAttributesRequest; import software.amazon.awssdk.services.s3.model.GetObjectAttributesResponse; import software.amazon.awssdk.services.s3.model.ObjectAttributes; import software.amazon.awssdk.services.s3.model.ObjectPart; import java.io.*; import java.nio.channels.FileChannel; import java.security.DigestInputStream; import java.security.MessageDigest; import java.security.NoSuchAlgorithmException; import java.util.Base64; import java.util.List; import static software.amazon.awssdk.services.s3.internal.resource.S3ResourceType.BUCKET; @Singleton public class S3ChecksumValidator { private final S3Client client; private final Logger logger; @Inject public S3ChecksumValidator(S3Client client, Logger logger) { this.client = client; this.logger = logger; } public boolean validateMultipartUpload(File file, String bucket, String s3Key) throws ChecksumValidationException { int chunkSize = 5 * 1024 * 1024; GetObjectAttributesResponse objectAttributes = client.getObjectAttributes(GetObjectAttributesRequest.builder().bucket(bucket).key(s3Key) .objectAttributes(ObjectAttributes.OBJECT_PARTS, ObjectAttributes.CHECKSUM).build()); try (InputStream localInput = new FileInputStream(file)) { MessageDigest sha256ChecksumOfChecksums = MessageDigest.getInstance("SHA-256"); MessageDigest sha256Part = MessageDigest.getInstance("SHA-256"); byte[] buffer = new byte[chunkSize]; int currentPart = 0; long partBreak = getPartBreak(objectAttributes, currentPart); int totalRead = 0; int read = localInput.read(buffer); while (read != -1) { totalRead += read; if (totalRead >= partBreak) { int difference = totalRead - (int) partBreak; byte[] partChecksum; if (totalRead != partBreak) { sha256Part.update(buffer, 0, read - difference); partChecksum = sha256Part.digest(); sha256ChecksumOfChecksums.update(partChecksum); sha256Part.reset(); sha256Part.update(buffer, read - difference, difference); } else { sha256Part.update(buffer, 0, read); partChecksum = sha256Part.digest(); sha256ChecksumOfChecksums.update(partChecksum); sha256Part.reset(); } String base64PartChecksum = Base64.getEncoder().encodeToString(partChecksum); if (!base64PartChecksum.equals(objectAttributes.objectParts().parts().get(currentPart).checksumSHA256())) { logger.info(String.format("Part checksum of local file does not match s3 file '%s'.", s3Key)); return false; } currentPart++; if (currentPart < objectAttributes.objectParts().totalPartsCount()) { partBreak += objectAttributes.objectParts().parts().get(currentPart - 1).size(); } } else { sha256Part.update(buffer, 0, read); } read = localInput.read(buffer); } logger.info(String.format("local parts: %s , remote parts: %s", currentPart + 1, objectAttributes.objectParts().totalPartsCount())); if (currentPart != objectAttributes.objectParts().totalPartsCount()) { currentPart++; byte[] partChecksum = sha256Part.digest(); sha256ChecksumOfChecksums.update(partChecksum); String base64PartChecksum = Base64.getEncoder().encodeToString(partChecksum); } String base64CalculatedChecksumOfChecksums = Base64.getEncoder().encodeToString(sha256ChecksumOfChecksums.digest()); if (!base64CalculatedChecksumOfChecksums.equals(objectAttributes.checksum().checksumSHA256())) { logger.info(String.format("Checksum of checksums of local file does not match s3 file '%s'.", s3Key)); logger.info(String.format("%s vs %s", base64CalculatedChecksumOfChecksums, objectAttributes.checksum().checksumSHA256())); return false; } } catch (IOException | NoSuchAlgorithmException e) { String msg = String.format("Could not read local checksum - %s", e.getMessage()); throw new ChecksumValidationException(msg, e); } return true; } private static long getPartBreak(GetObjectAttributesResponse objectAttributes, int currentPart) throws ChecksumValidationException { if(objectAttributes.objectParts() == null) { String msg = "Not a multipart upload - object attributes -> object parts is null"; throw new ChecksumValidationException(msg); } List<ObjectPart> parts = objectAttributes.objectParts().parts(); if(parts.isEmpty()) { String msg = "File was uploaded without checksum algorithm - object attributes -> object parts is empty"; throw new ChecksumValidationException(msg); } return parts.get(currentPart).size(); } }
问题
我先用一个约1.5GB的较小文件测试,代码运行正常。但在测试2.5GB的大文件时,函数返回false,日志输出如下:
[2024-02-01 18:34:08,768]-[Execution worker] INFO app.App - local parts: 125 , remote parts: 144 [2024-02-01 18:34:08,768]-[Execution worker] INFO app.App - Checksum of checksums of local file does not match s3 file 'shared/database.mmdb'. [2024-02-01 18:34:08,768]-[Execution worker] INFO app.App - faceAGZGc36kITYRStsK5zEw+iBJTgttwRWbmnQC+jQ= vs j3L01d+7qyiJ4zYSadr0/+N+Q8IfYbpWM7JTYvXrIlw=
看起来每个分片的校验和验证都通过了,但本地文件的分片数(125)远少于远程S3的分片数(144),导致校验和的校验和不匹配。我已确认本地和S3上的文件大小完全一致,请问可能是什么原因导致的?
问题原因
核心问题是**totalRead变量的整数溢出**:你用int类型存储已读取的总字节数,而int的最大值是2^31-1(约2.1GB)。当读取2.5GB的大文件时,totalRead会溢出变为负数,导致后续totalRead >= partBreak的判断逻辑完全失效,无法正确分割本地文件的分片,最终统计出的本地分片数远小于S3的实际分片数。
此外还有一个次要逻辑问题:partBreak的更新逻辑错误。当前循环中partBreak += objectAttributes.objectParts().parts().get(currentPart - 1).size()是累加前一个分片的大小,而正确的逻辑应该是累加当前分片的大小,不过这个问题在文件小于2.1GB时不会暴露。
修复方案
- 修改
totalRead的类型为long:避免大文件读取时的整数溢出:long totalRead = 0; - 修正
difference的计算与类型转换:long difference = totalRead - partBreak; // 因为单个分片大小不会超过int范围,所以可以安全转换 int diffInt = (int) difference; - 修复
partBreak的更新逻辑:if (currentPart < objectAttributes.objectParts().totalPartsCount()) { partBreak += objectAttributes.objectParts().parts().get(currentPart).size(); }
完成这些修改后,大文件的分片分割逻辑会与S3的实际分片一致,校验和的校验和也能正确匹配。
内容的提问来源于stack exchange,提问作者Nonetallt

