SpringBoot应用复制S3文件时IOUtils.copyLarge引发OOM及411问题求解
我有一个部署在Kubernetes中的SpringBoot API应用,单实例配置为cpu:200m、memory:8Gi,JVM参数设置为_JAVA_OPTIONS: -Xmx7G -Xms7G。需求是通过预签名URL在两个S3存储桶之间复制大小介于30Mb至1Gb的文件。
当用1Gb文件并发调用该API超过2次时,会触发java.lang.OutOfMemoryError: Java heap space错误。
以下是我的代码:
public ResponseEntity<String> transferData(Config config) throws IOException { var inputUrl = config.getInputdUrl(); var outputUrl = config.getOutputdUrl(); var uploadConnection = uploader.getUploadConnection(outputUrl); try (var inputStream = downloader.getDownloadConnection(inputUrl); var outputStream = uploadConnection.getOutputStream()) { IOUtils.copyLarge(inputStream, outputStream); var resp = uploadConnection.getResponseCode(); return ResponseEntity .status(HttpStatus.OK) .body("Transfer done with response:" + resp); } catch (IOException ioe) { return ResponseEntity .status(HttpStatus.INTERNAL_SERVER_ERROR) .body("An error occurred while transferring data : " + ioe); } }
getUploadConnection返回HttpURLConnection:
var connection = (HttpURLConnection) new URL(presignedUrl).openConnection(); connection.setDoOutput(true); connection.setRequestMethod(PUT.toString()); return connection;
getDownloadConnection返回HttpURLConnection的输入流:
HttpURLConnection connection = (HttpURLConnection) new URL(presignedUrl).openConnection(); connection.setRequestMethod(GET.toString()); connection.setDoOutput(true); return connection.getInputStream();
我原本认为流式处理不会将文件完全加载到内存,对内存的影响应该不大,但实际出现了问题,请问我忽略了什么?
编辑1:
我在getUploadConnection中添加了setChunkedStreamingMode,现在该方法返回:
var connection = (HttpURLConnection) new URL(url).openConnection(); connection.setDoOutput(true); connection.setDoInput(true); connection.setChunkedStreamingMode(-1); connection.setRequestMethod(PUT.toString()); return connection;
但此时收到了S3存储桶返回的411错误。
1. 为什么会触发OOM?
你误以为是纯流式处理,但HttpURLConnection默认会将整个请求体缓存到内存中,直到所有数据准备完毕才会发送给S3。当处理1GB文件时,每个并发请求都会在堆内存中缓存完整的1GB数据,2次并发就会占用至少2GB内存,再加上应用本身的内存开销,很容易触发堆内存溢出。
另外,getDownloadConnection里的connection.setDoOutput(true)是无效配置——GET请求不需要输出流,虽然不会直接导致OOM,但属于冗余代码。
2. 为什么添加setChunkedStreamingMode(-1)会返回411错误?
S3不支持PUT请求使用分块传输编码(Chunked Transfer Encoding)。setChunkedStreamingMode(-1)会让请求采用分块编码格式,而S3要求PUT请求必须携带Content-Length头明确文件大小,否则就会返回411(Length Required)错误。
3. 正确的解决方法
方案一:使用S3原生跨桶复制(推荐)
两个S3桶之间的文件复制,完全不需要通过你的应用中转数据。直接调用S3的CopyObject API(或对应SDK方法),让S3内部完成复制,你的应用只需要发起复制请求即可,完全不占用应用的内存和带宽。
示例代码(AWS SDK for Java):
// 初始化S3客户端 AmazonS3 s3Client = AmazonS3ClientBuilder.defaultClient(); // 构造跨桶复制请求 CopyObjectRequest copyRequest = new CopyObjectRequest( "源存储桶名称", "源对象键", "目标存储桶名称", "目标对象键" ); // 执行复制 s3Client.copyObject(copyRequest);
如果需要用预签名URL的场景,也可以通过SDK生成预签名的复制请求,全程无需处理文件流。
方案二:必须通过应用中转时的处理(不推荐)
如果受限于业务规则必须通过应用中转,需要让HttpURLConnection不缓存完整请求体,同时提供S3要求的Content-Length头:
- 修改下载逻辑,获取源文件大小:
// 新增辅助类存储输入流和文件大小 public static class DownloadInfo { private final InputStream inputStream; private final long contentLength; public DownloadInfo(InputStream inputStream, long contentLength) { this.inputStream = inputStream; this.contentLength = contentLength; } // getter方法省略 } // 修改getDownloadConnection方法 public DownloadInfo getDownloadConnection(String presignedUrl) throws IOException { HttpURLConnection connection = (HttpURLConnection) new URL(presignedUrl).openConnection(); connection.setRequestMethod("GET"); // 移除无效的setDoOutput(true) long contentLength = connection.getContentLengthLong(); InputStream inputStream = connection.getInputStream(); return new DownloadInfo(inputStream, contentLength); }
- 修改上传逻辑,设置固定长度流模式:
public HttpURLConnection getUploadConnection(String presignedUrl, long contentLength) throws IOException { HttpURLConnection connection = (HttpURLConnection) new URL(presignedUrl).openConnection(); connection.setDoOutput(true); connection.setRequestMethod("PUT"); // 设置Content-Length,满足S3要求 connection.setFixedLengthStreamingMode(contentLength); // 禁用请求缓存,避免内存堆积 connection.setUseCaches(false); return connection; }
- 修改传输逻辑,用固定缓冲区复制流:
public ResponseEntity<String> transferData(Config config) throws IOException { var inputUrl = config.getInputdUrl(); var outputUrl = config.getOutputdUrl(); var downloadInfo = downloader.getDownloadConnection(inputUrl); var uploadConnection = uploader.getUploadConnection(outputUrl, downloadInfo.getContentLength()); try (var inputStream = downloadInfo.getInputStream(); var outputStream = uploadConnection.getOutputStream()) { // 使用8KB固定缓冲区,避免内存占用过大 byte[] buffer = new byte[8192]; int bytesRead; while ((bytesRead = inputStream.read(buffer)) != -1) { outputStream.write(buffer, 0, bytesRead); // 强制刷新输出流,避免缓冲区堆积 outputStream.flush(); } var resp = uploadConnection.getResponseCode(); return ResponseEntity .status(HttpStatus.OK) .body("Transfer done with response:" + resp); } catch (IOException ioe) { return ResponseEntity .status(HttpStatus.INTERNAL_SERVER_ERROR) .body("An error occurred while transferring data : " + ioe); } }
这样设置后,HttpURLConnection会按照固定长度分块发送数据,不会将整个文件缓存到内存中,同时S3能获取到文件大小,不会返回411错误。
额外优化建议
- 调整JVM参数:
-Xmx7G设置过大,Kubernetes分配的8Gi内存需要预留容器系统进程和JVM非堆内存的空间,建议改为-Xmx6G,避免触发Kubernetes的OOM Killer。 - 限制API并发数:通过Spring的限流注解(如
@RateLimiter)或Kubernetes的限流策略,避免同时处理过多大文件传输请求。
内容的提问来源于stack exchange,提问作者ismala

