You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SpringBoot应用复制S3文件时IOUtils.copyLarge引发OOM及411问题求解

问题描述

我有一个部署在Kubernetes中的SpringBoot API应用,单实例配置为cpu:200m、memory:8Gi,JVM参数设置为_JAVA_OPTIONS: -Xmx7G -Xms7G。需求是通过预签名URL在两个S3存储桶之间复制大小介于30Mb至1Gb的文件。

当用1Gb文件并发调用该API超过2次时,会触发java.lang.OutOfMemoryError: Java heap space错误。

以下是我的代码:

public ResponseEntity<String> transferData(Config config) throws IOException {

    var inputUrl = config.getInputdUrl();
    var outputUrl = config.getOutputdUrl();


    var uploadConnection = uploader.getUploadConnection(outputUrl);

    try (var inputStream = downloader.getDownloadConnection(inputUrl); 
         var outputStream = uploadConnection.getOutputStream()) {

        IOUtils.copyLarge(inputStream, outputStream);

        var resp = uploadConnection.getResponseCode();

        return ResponseEntity
                .status(HttpStatus.OK)
                .body("Transfer done with response:" + resp);
    } catch (IOException ioe) {
        return ResponseEntity
                .status(HttpStatus.INTERNAL_SERVER_ERROR)
                .body("An error occurred while transferring data : " + ioe);
    }

}

getUploadConnection返回HttpURLConnection:

var connection = (HttpURLConnection) new URL(presignedUrl).openConnection();
connection.setDoOutput(true);
connection.setRequestMethod(PUT.toString());
return connection;

getDownloadConnection返回HttpURLConnection的输入流:

HttpURLConnection connection = (HttpURLConnection) new URL(presignedUrl).openConnection();
connection.setRequestMethod(GET.toString());
connection.setDoOutput(true);
return connection.getInputStream();

我原本认为流式处理不会将文件完全加载到内存,对内存的影响应该不大,但实际出现了问题,请问我忽略了什么?

编辑1:
我在getUploadConnection中添加了setChunkedStreamingMode,现在该方法返回:

var connection = (HttpURLConnection) new URL(url).openConnection();
connection.setDoOutput(true);
connection.setDoInput(true);
connection.setChunkedStreamingMode(-1);
connection.setRequestMethod(PUT.toString());
return connection;

但此时收到了S3存储桶返回的411错误。


问题分析与解决方案

1. 为什么会触发OOM?

你误以为是纯流式处理,但HttpURLConnection默认会将整个请求体缓存到内存中,直到所有数据准备完毕才会发送给S3。当处理1GB文件时,每个并发请求都会在堆内存中缓存完整的1GB数据,2次并发就会占用至少2GB内存,再加上应用本身的内存开销,很容易触发堆内存溢出。

另外,getDownloadConnection里的connection.setDoOutput(true)是无效配置——GET请求不需要输出流,虽然不会直接导致OOM,但属于冗余代码。

2. 为什么添加setChunkedStreamingMode(-1)会返回411错误?

S3不支持PUT请求使用分块传输编码(Chunked Transfer Encoding)。setChunkedStreamingMode(-1)会让请求采用分块编码格式,而S3要求PUT请求必须携带Content-Length头明确文件大小,否则就会返回411(Length Required)错误。

3. 正确的解决方法

方案一:使用S3原生跨桶复制(推荐)

两个S3桶之间的文件复制,完全不需要通过你的应用中转数据。直接调用S3的CopyObject API(或对应SDK方法),让S3内部完成复制,你的应用只需要发起复制请求即可,完全不占用应用的内存和带宽。

示例代码(AWS SDK for Java):

// 初始化S3客户端
AmazonS3 s3Client = AmazonS3ClientBuilder.defaultClient();

// 构造跨桶复制请求
CopyObjectRequest copyRequest = new CopyObjectRequest(
    "源存储桶名称", "源对象键",
    "目标存储桶名称", "目标对象键"
);

// 执行复制
s3Client.copyObject(copyRequest);

如果需要用预签名URL的场景,也可以通过SDK生成预签名的复制请求,全程无需处理文件流。

方案二:必须通过应用中转时的处理(不推荐)

如果受限于业务规则必须通过应用中转,需要让HttpURLConnection不缓存完整请求体,同时提供S3要求的Content-Length头:

  1. 修改下载逻辑,获取源文件大小:
// 新增辅助类存储输入流和文件大小
public static class DownloadInfo {
    private final InputStream inputStream;
    private final long contentLength;

    public DownloadInfo(InputStream inputStream, long contentLength) {
        this.inputStream = inputStream;
        this.contentLength = contentLength;
    }

    // getter方法省略
}

// 修改getDownloadConnection方法
public DownloadInfo getDownloadConnection(String presignedUrl) throws IOException {
    HttpURLConnection connection = (HttpURLConnection) new URL(presignedUrl).openConnection();
    connection.setRequestMethod("GET");
    // 移除无效的setDoOutput(true)
    long contentLength = connection.getContentLengthLong();
    InputStream inputStream = connection.getInputStream();
    return new DownloadInfo(inputStream, contentLength);
}
  1. 修改上传逻辑,设置固定长度流模式:
public HttpURLConnection getUploadConnection(String presignedUrl, long contentLength) throws IOException {
    HttpURLConnection connection = (HttpURLConnection) new URL(presignedUrl).openConnection();
    connection.setDoOutput(true);
    connection.setRequestMethod("PUT");
    // 设置Content-Length,满足S3要求
    connection.setFixedLengthStreamingMode(contentLength);
    // 禁用请求缓存,避免内存堆积
    connection.setUseCaches(false);
    return connection;
}
  1. 修改传输逻辑,用固定缓冲区复制流:
public ResponseEntity<String> transferData(Config config) throws IOException {
    var inputUrl = config.getInputdUrl();
    var outputUrl = config.getOutputdUrl();

    var downloadInfo = downloader.getDownloadConnection(inputUrl);
    var uploadConnection = uploader.getUploadConnection(outputUrl, downloadInfo.getContentLength());

    try (var inputStream = downloadInfo.getInputStream(); 
         var outputStream = uploadConnection.getOutputStream()) {
        // 使用8KB固定缓冲区,避免内存占用过大
        byte[] buffer = new byte[8192];
        int bytesRead;
        while ((bytesRead = inputStream.read(buffer)) != -1) {
            outputStream.write(buffer, 0, bytesRead);
            // 强制刷新输出流,避免缓冲区堆积
            outputStream.flush();
        }

        var resp = uploadConnection.getResponseCode();
        return ResponseEntity
                .status(HttpStatus.OK)
                .body("Transfer done with response:" + resp);
    } catch (IOException ioe) {
        return ResponseEntity
                .status(HttpStatus.INTERNAL_SERVER_ERROR)
                .body("An error occurred while transferring data : " + ioe);
    }
}

这样设置后,HttpURLConnection会按照固定长度分块发送数据,不会将整个文件缓存到内存中,同时S3能获取到文件大小,不会返回411错误。

额外优化建议

  • 调整JVM参数:-Xmx7G设置过大,Kubernetes分配的8Gi内存需要预留容器系统进程和JVM非堆内存的空间,建议改为-Xmx6G,避免触发Kubernetes的OOM Killer。
  • 限制API并发数:通过Spring的限流注解(如@RateLimiter)或Kubernetes的限流策略,避免同时处理过多大文件传输请求。

内容的提问来源于stack exchange,提问作者ismala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 00:00:37