使用Java SDK v2批量拷贝AWS S3文件时如何避免偶发的服务端响应异常
S3高并发数据迁移偶发连接异常问题
我在数据迁移过程中需要拷贝数百万个S3文件,计划执行高并发拷贝操作,当前使用Java SDK v2 API,但在批量并发拷贝S3文件时频繁出现罕见的偶发异常。
常见异常类型
最常遇到的异常如下:
Unable to execute HTTP request: Server failed to send complete response. The channel was closed. This may have been done by the client (e.g. because the request was aborted), by the service (e.g. because there was a handshake error, the request took too long, or the client tried to write on a read-only socket), or by an intermediary party (e.g. because the channel was idle for too long).
其余遇到的异常如下:
Unable to execute HTTP request: The channel was closed. This may have been done by the client (e.g. because the request was aborted), by the service (e.g. because there was a handshake error, the request took too long, or the client tried to write on a read-only socket), or by an intermediary party (e.g. because the channel was idle for too long).
We encountered an internal error. Please try again. (Service: S3, Status Code: 500 ...
Unable to execute HTTP request: Channel was closed before it could be written to.
问题复现代码
以下是大多情况下可稳定触发该问题的示例代码(该问题受异常/繁忙的S3服务节点、网络流量、限流或竞态条件等因素影响,无法100%稳定复现):
NettyNioAsyncHttpClient.Builder asyncHttpClientBuilder = NettyNioAsyncHttpClient.builder() .maxConcurrency(300) .maxPendingConnectionAcquires(500) .connectionTimeout(Duration.ofSeconds(10)) .connectionAcquisitionTimeout(Duration.ofSeconds(60)) .readTimeout(Duration.ofSeconds(60)); S3AsyncClient s3Client = S3AsyncClient.builder() .credentialsProvider(awsCredentialsProvider) .region(Region.US_EAST_1) // 即便配置重试也会出现相同失败,只是触发时间更晚 .overrideConfiguration(config -> config.retryPolicy(RetryPolicy.none()).build()) .httpClientBuilder(asyncHttpClientBuilder) .build(); List<CompletableFuture> futures = new ArrayList<>(); for (int i = 1; i <= 500; i++) { String key = "zerotestfile"; Path outFile = Paths.get("/tmp/experiment/").resolve(key + "-" + i); outFile.getParent().toFile().mkdirs(); if (outFile.toFile().exists()) { outFile.toFile().delete(); } log.info("Downloading: {} ({})", key, i); GetObjectRequest request = GetObjectRequest.builder() .bucket("my-test-bucket") .key(key) .build(); CompletableFuture<GetObjectResponse> future = s3Client.getObject(request, AsyncResponseTransformer.toFile(outFile)) .exceptionally(exception -> { log.error("Error", exception); return null; }); futures.add(future); } CompletableFuture.allOf(futures.toArray(new CompletableFuture[0])).join();
测试使用的500k文件生成命令为:dd if=/dev/zero of=zerotestfile bs=1024 count=500
问题说明
使用默认重试策略似乎可以修复上述示例中的问题,但在实际迁移数百万个文件的场景下,即便配置了重试策略,依然会复现上述所有异常。
补充说明:实际迁移逻辑使用跨区域CopyObject调用,为简化问题,示例改为单区域GetObject请求,将maxConcurrency设为2000、执行2500次拷贝操作时也会出现类似错误。
已解决的异常及对应配置
通过配置调整已修复如下错误:
- 错误:
Unable to execute HTTP request: Acquire operation took longer than the configured maximum time.
解决配置:.connectionAcquisitionTimeout(Duration.ofSeconds(60)) - 错误:
Unable to execute HTTP request: connection timed out
解决配置:.connectionTimeout(Duration.ofSeconds(10)) - 错误:
ReadTimeoutException: null
解决配置:.readTimeout(Duration.ofSeconds(60)) - 错误:
Acquire operation took longer than the configured maximum time. This indicates that a request cannot get a connection from the pool within the specified maximum time. This can be due to high request rate. Consider taking any of the following actions to mitigate the issue: increase max connections, increase acquire timeout, or slowing the request rate.
解决配置:.maxPendingConnectionAcquires(500)
内容的提问来源于stack exchange,提问作者Curtis Yallop
相关产品推荐
相关产品推荐

