You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java 8下Multipart大附件流提取失败(堆内存溢出)求助

Multipart大文件附件提取堆内存溢出问题排查(Java 8)

业务场景:使用Java 8编写代码接收Multipart格式输入载荷,提取最大2GB的附件,必须以流式方式处理。当前代码处理小文件正常,但处理大文件时触发堆内存错误,流式处理未生效。

输入示例

multipart/related; boundary="boundary123"
--boundary123
Content-Type: text/xml; charset=UTF-8
Content-Transfer-Encoding: 8bit

<soapenv:Envelope xmlns:soapenv="http://schemas.xmlsoap.org/soap/envelope/">
  <soapenv:Body>
    <ns:getUserDetailsResponse xmlns:ns="http://example.com">
      <ns:return>
        <name>John Doe</name>
        <age>42</age>
      </ns:return>
    </ns:getUserDetailsResponse>
  </soapenv:Body>
</soapenv:Envelope>
--boundary123
Content-Type: application/octet-stream
Content-Transfer-Encoding: binary
Content-Disposition: attachment; filename="output.txt"

YWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWF==
--boundary123--

预期输出

YWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWF==

现有代码

import java.io.BufferedReader;
import java.io.ByteArrayInputStream;
import java.io.IOException;
import java.io.InputStream;
import java.io.InputStreamReader;
import java.io.PipedInputStream;
import java.io.PipedOutputStream;
import java.nio.charset.StandardCharsets;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

import javax.activation.DataSource;
import javax.mail.MessagingException;
import javax.mail.util.ByteArrayDataSource;
import javax.mail.internet.MimeBodyPart;
import javax.mail.internet.MimeMultipart;

public class AttachmentExtractor {
    public static InputStream extract(String message, int bufferSize) throws IOException, MessagingException {

        InputStream input = new ByteArrayInputStream(message.getBytes(StandardCharsets.UTF_8));

        PipedOutputStream attachmentOutStream = new PipedOutputStream();
        PipedInputStream pipedStream = new PipedInputStream(attachmentOutStream);
        
        Thread newThread = new Thread(() -> {
            
        //  int chunkCount = 1 ;    /* only used for testing */
            
            int bytesRead;
            String boundary = null;
            boolean isFirstChunk = true;
            boolean isLastChunk = false;
            
            byte[] buffer = new byte[bufferSize];
            
            try {
                
                if(input.available() < 3 * bufferSize)
                    buffer = new byte[input.available()];
                
                while ((bytesRead = input.read(buffer)) != -1) {
                    
                    // System.out.println("Chunk #"+chunkCount);  /* only used for testing */

                    if(isFirstChunk) {
                        
                        String firstChunk = new String(buffer, 0, bytesRead, StandardCharsets.UTF_8);
                        
                        Pattern pattern = Pattern.compile("boundary=\"(.*)\"");
                        Matcher matcher = pattern.matcher(firstChunk);
                        if (matcher.find()) {
                            boundary = matcher.group(1);
                        } else {
                            throw new IllegalArgumentException("Unable to extract boundary from input stream.");
                        }

                        DataSource bufferDS = new ByteArrayDataSource(firstChunk.getBytes(StandardCharsets.UTF_8), "multipart/form-data");  
                        MimeMultipart multipart = new MimeMultipart(bufferDS);
                        int attachmentPartIndex = multipart.getCount() - 1;
                        MimeBodyPart attachmentPart = (MimeBodyPart) multipart.getBodyPart(attachmentPartIndex);

                        attachmentPart.getDataHandler().writeTo(attachmentOutStream);

                        isFirstChunk = false;
                        
                    }else if (isLastChunk) {
                        
                        InputStream bufferIS = new ByteArrayInputStream(buffer);
                        BufferedReader reader = new BufferedReader(new InputStreamReader(bufferIS));
                        String line;
                        while(reader.ready() && (line=reader.readLine()) != null) {
                            
                            if(line.indexOf(boundary)>0)
                                continue;
                            
                            attachmentOutStream.write(line.getBytes(StandardCharsets.UTF_8));
                        }
                        
                    }else
                        attachmentOutStream.write(buffer);
                    
                    if(input.available() <= bufferSize)
                        isLastChunk = true;
                    
                    // chunkCount++;  /* only used for testing */
                }
                
                try {
                    attachmentOutStream.close();
                } catch (IOException e) { /* Output Stream already closed - can be ignored. */ }
                
            }catch(Exception e) {
                e.printStackTrace();
            }
        });
        
        newThread.start();

        return pipedStream;        

    }
}

问题根源分析

  1. 全量加载输入到内存:方法参数为String message,意味着整个Multipart内容已被读入内存转为字符串,大文件直接导致堆内存占用过高。
  2. 内存型IO类误用:ByteArrayInputStream和ByteArrayDataSource都是基于内存数组的实现,会将输入内容全部加载到内存,完全违背流式处理的初衷。
  3. input.available()依赖错误:该方法仅返回当前可无阻塞读取的字节数,无法判断剩余总字节数,对于大文件此值不可靠,导致isLastChunk判断错误,处理逻辑混乱。
  4. Multipart解析不完整:仅用第一块内容初始化MimeMultipart,无法处理跨块的Multipart结构,大文件附件内容必然跨块,导致后续直接写入整个buffer(包含边界等无关内容),同时第一块处理时可能重复写入部分附件内容。

修复方案(Java 8兼容)

核心调整点

  • 直接接收InputStream作为输入,避免全量加载到内存。
  • 手动实现流式边界匹配与内容提取,放弃依赖内存型Multipart解析类。
  • 移除input.available()的依赖,改为逐行扫描边界。

修复后的代码示例

import java.io.*;
import java.nio.charset.StandardCharsets;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class StreamingAttachmentExtractor {

    private static final Pattern BOUNDARY_PATTERN = Pattern.compile("boundary=\"(.*)\"");
    private static final String CONTENT_DISPOSITION_ATTACHMENT = "Content-Disposition: attachment";

    public static InputStream extract(InputStream inputStream) throws IOException {
        PipedOutputStream output = new PipedOutputStream();
        PipedInputStream result = new PipedInputStream(output);

        new Thread(() -> {
            try (BufferedReader reader = new BufferedReader(new InputStreamReader(inputStream, StandardCharsets.UTF_8));
                 BufferedWriter writer = new BufferedWriter(new OutputStreamWriter(output, StandardCharsets.UTF_8))) {

                String boundary = extractBoundary(reader);
                boolean inAttachment = false;
                String line;

                while ((line = reader.readLine()) != null) {
                    // 检查是否进入新的Multipart块
                    if (line.startsWith("--" + boundary)) {
                        inAttachment = false;
                        // 读取块头信息,直到空行
                        while ((line = reader.readLine()) != null && !line.isEmpty()) {
                            if (line.contains(CONTENT_DISPOSITION_ATTACHMENT)) {
                                inAttachment = true;
                            }
                        }
                        continue;
                    }
                    // 检查是否到达Multipart结束标记
                    if (line.equals("--" + boundary + "--")) {
                        break;
                    }
                    // 处于附件部分时写入内容
                    if (inAttachment) {
                        writer.write(line);
                        writer.newLine();
                    }
                }
                writer.flush();
            } catch (IOException e) {
                e.printStackTrace();
            } finally {
                try {
                    output.close();
                } catch (IOException ignored) {}
            }
        }).start();

        return result;
    }

    private static String extractBoundary(BufferedReader reader) throws IOException {
        String line;
        while ((line = reader.readLine()) != null) {
            if (line.startsWith("multipart/")) {
                Matcher matcher = BOUNDARY_PATTERN.matcher(line);
                if (matcher.find()) {
                    return matcher.group(1);
                }
                throw new IllegalArgumentException("No boundary found in multipart header");
            }
        }
        throw new IllegalArgumentException("Invalid multipart input");
    }
}

代码说明

  1. 流式输入处理:直接接收InputStream,通过BufferedReader逐行读取,内存仅保留当前行数据,避免全量加载。
  2. 边界提取逻辑:从Multipart头行中提取boundary,无需加载全部内容。
  3. 附件识别与写入:逐行扫描,遇到附件的Content-Disposition头后开始写入内容,直到遇到结束边界,全程流式处理。
  4. 异步管道流:使用PipedInputStream和PipedOutputStream实现异步流式输出,避免阻塞调用线程。

使用示例

public class Main {
    public static void main(String[] args) throws IOException {
        try (InputStream input = new FileInputStream("large-multipart.input");
             InputStream attachment = StreamingAttachmentExtractor.extract(input);
             OutputStream output = new FileOutputStream("attachment.output")) {
            byte[] buffer = new byte[8192];
            int bytesRead;
            while ((bytesRead = attachment.read(buffer)) != -1) {
                output.write(buffer, 0, bytesRead);
            }
        }
    }
}

内容的提问来源于stack exchange,提问作者Arpit Goel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 22:46:59