Java 8下Multipart大附件流提取失败(堆内存溢出)求助
Multipart大文件附件提取堆内存溢出问题排查(Java 8)
业务场景:使用Java 8编写代码接收Multipart格式输入载荷,提取最大2GB的附件,必须以流式方式处理。当前代码处理小文件正常,但处理大文件时触发堆内存错误,流式处理未生效。
输入示例
multipart/related; boundary="boundary123" --boundary123 Content-Type: text/xml; charset=UTF-8 Content-Transfer-Encoding: 8bit <soapenv:Envelope xmlns:soapenv="http://schemas.xmlsoap.org/soap/envelope/"> <soapenv:Body> <ns:getUserDetailsResponse xmlns:ns="http://example.com"> <ns:return> <name>John Doe</name> <age>42</age> </ns:return> </ns:getUserDetailsResponse> </soapenv:Body> </soapenv:Envelope> --boundary123 Content-Type: application/octet-stream Content-Transfer-Encoding: binary Content-Disposition: attachment; filename="output.txt" YWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWF== --boundary123--
预期输出
YWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWFhYWF==
现有代码
import java.io.BufferedReader; import java.io.ByteArrayInputStream; import java.io.IOException; import java.io.InputStream; import java.io.InputStreamReader; import java.io.PipedInputStream; import java.io.PipedOutputStream; import java.nio.charset.StandardCharsets; import java.util.regex.Matcher; import java.util.regex.Pattern; import javax.activation.DataSource; import javax.mail.MessagingException; import javax.mail.util.ByteArrayDataSource; import javax.mail.internet.MimeBodyPart; import javax.mail.internet.MimeMultipart; public class AttachmentExtractor { public static InputStream extract(String message, int bufferSize) throws IOException, MessagingException { InputStream input = new ByteArrayInputStream(message.getBytes(StandardCharsets.UTF_8)); PipedOutputStream attachmentOutStream = new PipedOutputStream(); PipedInputStream pipedStream = new PipedInputStream(attachmentOutStream); Thread newThread = new Thread(() -> { // int chunkCount = 1 ; /* only used for testing */ int bytesRead; String boundary = null; boolean isFirstChunk = true; boolean isLastChunk = false; byte[] buffer = new byte[bufferSize]; try { if(input.available() < 3 * bufferSize) buffer = new byte[input.available()]; while ((bytesRead = input.read(buffer)) != -1) { // System.out.println("Chunk #"+chunkCount); /* only used for testing */ if(isFirstChunk) { String firstChunk = new String(buffer, 0, bytesRead, StandardCharsets.UTF_8); Pattern pattern = Pattern.compile("boundary=\"(.*)\""); Matcher matcher = pattern.matcher(firstChunk); if (matcher.find()) { boundary = matcher.group(1); } else { throw new IllegalArgumentException("Unable to extract boundary from input stream."); } DataSource bufferDS = new ByteArrayDataSource(firstChunk.getBytes(StandardCharsets.UTF_8), "multipart/form-data"); MimeMultipart multipart = new MimeMultipart(bufferDS); int attachmentPartIndex = multipart.getCount() - 1; MimeBodyPart attachmentPart = (MimeBodyPart) multipart.getBodyPart(attachmentPartIndex); attachmentPart.getDataHandler().writeTo(attachmentOutStream); isFirstChunk = false; }else if (isLastChunk) { InputStream bufferIS = new ByteArrayInputStream(buffer); BufferedReader reader = new BufferedReader(new InputStreamReader(bufferIS)); String line; while(reader.ready() && (line=reader.readLine()) != null) { if(line.indexOf(boundary)>0) continue; attachmentOutStream.write(line.getBytes(StandardCharsets.UTF_8)); } }else attachmentOutStream.write(buffer); if(input.available() <= bufferSize) isLastChunk = true; // chunkCount++; /* only used for testing */ } try { attachmentOutStream.close(); } catch (IOException e) { /* Output Stream already closed - can be ignored. */ } }catch(Exception e) { e.printStackTrace(); } }); newThread.start(); return pipedStream; } }
问题根源分析
- 全量加载输入到内存:方法参数为
String message,意味着整个Multipart内容已被读入内存转为字符串,大文件直接导致堆内存占用过高。 - 内存型IO类误用:
ByteArrayInputStream和ByteArrayDataSource都是基于内存数组的实现,会将输入内容全部加载到内存,完全违背流式处理的初衷。 input.available()依赖错误:该方法仅返回当前可无阻塞读取的字节数,无法判断剩余总字节数,对于大文件此值不可靠,导致isLastChunk判断错误,处理逻辑混乱。- Multipart解析不完整:仅用第一块内容初始化
MimeMultipart,无法处理跨块的Multipart结构,大文件附件内容必然跨块,导致后续直接写入整个buffer(包含边界等无关内容),同时第一块处理时可能重复写入部分附件内容。
修复方案(Java 8兼容)
核心调整点
- 直接接收
InputStream作为输入,避免全量加载到内存。 - 手动实现流式边界匹配与内容提取,放弃依赖内存型Multipart解析类。
- 移除
input.available()的依赖,改为逐行扫描边界。
修复后的代码示例
import java.io.*; import java.nio.charset.StandardCharsets; import java.util.regex.Matcher; import java.util.regex.Pattern; public class StreamingAttachmentExtractor { private static final Pattern BOUNDARY_PATTERN = Pattern.compile("boundary=\"(.*)\""); private static final String CONTENT_DISPOSITION_ATTACHMENT = "Content-Disposition: attachment"; public static InputStream extract(InputStream inputStream) throws IOException { PipedOutputStream output = new PipedOutputStream(); PipedInputStream result = new PipedInputStream(output); new Thread(() -> { try (BufferedReader reader = new BufferedReader(new InputStreamReader(inputStream, StandardCharsets.UTF_8)); BufferedWriter writer = new BufferedWriter(new OutputStreamWriter(output, StandardCharsets.UTF_8))) { String boundary = extractBoundary(reader); boolean inAttachment = false; String line; while ((line = reader.readLine()) != null) { // 检查是否进入新的Multipart块 if (line.startsWith("--" + boundary)) { inAttachment = false; // 读取块头信息,直到空行 while ((line = reader.readLine()) != null && !line.isEmpty()) { if (line.contains(CONTENT_DISPOSITION_ATTACHMENT)) { inAttachment = true; } } continue; } // 检查是否到达Multipart结束标记 if (line.equals("--" + boundary + "--")) { break; } // 处于附件部分时写入内容 if (inAttachment) { writer.write(line); writer.newLine(); } } writer.flush(); } catch (IOException e) { e.printStackTrace(); } finally { try { output.close(); } catch (IOException ignored) {} } }).start(); return result; } private static String extractBoundary(BufferedReader reader) throws IOException { String line; while ((line = reader.readLine()) != null) { if (line.startsWith("multipart/")) { Matcher matcher = BOUNDARY_PATTERN.matcher(line); if (matcher.find()) { return matcher.group(1); } throw new IllegalArgumentException("No boundary found in multipart header"); } } throw new IllegalArgumentException("Invalid multipart input"); } }
代码说明
- 流式输入处理:直接接收
InputStream,通过BufferedReader逐行读取,内存仅保留当前行数据,避免全量加载。 - 边界提取逻辑:从Multipart头行中提取boundary,无需加载全部内容。
- 附件识别与写入:逐行扫描,遇到附件的
Content-Disposition头后开始写入内容,直到遇到结束边界,全程流式处理。 - 异步管道流:使用
PipedInputStream和PipedOutputStream实现异步流式输出,避免阻塞调用线程。
使用示例
public class Main { public static void main(String[] args) throws IOException { try (InputStream input = new FileInputStream("large-multipart.input"); InputStream attachment = StreamingAttachmentExtractor.extract(input); OutputStream output = new FileOutputStream("attachment.output")) { byte[] buffer = new byte[8192]; int bytesRead; while ((bytesRead = attachment.read(buffer)) != -1) { output.write(buffer, 0, bytesRead); } } } }
内容的提问来源于stack exchange,提问作者Arpit Goel
相关产品推荐
相关产品推荐

