Java处理大XML移除CDATA时出现堆内存溢出该如何解决
解决方案
核心思路
不要按行读取,采用固定大小的字符缓冲区做流式逐块处理,全程仅保留缓冲区大小的内存占用,无需加载整个文件或超长行到堆内存,同时处理标记跨缓冲区的匹配问题,完全不会触发OOM。
具体实现方案
方案1:全局替换CDATA标记(保留外层XML结构)
适合需要保留原外层XML结构,仅移除CDATA标记的场景:
- 设定缓冲区大小(建议4096/8192字符,内存占用仅几KB)
- 维护一个长度为标记最大长度减1的剩余缓存,用来存储上一次读取块的末尾字符,解决CDATA标记跨块拆分的问题
- 逐块读取内容,拼接剩余缓存后匹配
<![CDATA[和]]>标记,匹配到标记就跳过写入,其余内容原封不动写入输出流
private void replaceCdataByStream(Reader reader, String outputPath) throws IOException { final int BUFFER_SIZE = 8192; // CDATA两个标记的最大长度是9(<![CDATA[长度9) final int MAX_MARKER_LEN = 9; char[] buffer = new char[BUFFER_SIZE]; char[] leftover = new char[MAX_MARKER_LEN - 1]; int leftoverLen = 0; try (BufferedWriter writer = new BufferedWriter(new FileWriter(outputPath))) { int readLen; while ((readLen = reader.read(buffer)) != -1) { // 拼接上一次剩余的字符和本次读取的内容 char[] fullBuf = new char[leftoverLen + readLen]; System.arraycopy(leftover, 0, fullBuf, 0, leftoverLen); System.arraycopy(buffer, 0, fullBuf, leftoverLen, readLen); int pos = 0; int currentLen = fullBuf.length; while (pos <= currentLen - MAX_MARKER_LEN) { // 匹配<![CDATA[ if (fullBuf[pos] == '<' && pos + 8 < currentLen && fullBuf[pos+1] == '!' && fullBuf[pos+2] == '[' && fullBuf[pos+3] == 'C' && fullBuf[pos+4] == 'D' && fullBuf[pos+5] == 'A' && fullBuf[pos+6] == 'T' && fullBuf[pos+7] == 'A' && fullBuf[pos+8] == '[') { writer.write(fullBuf, 0, pos); pos += 9; // 跳过标记后,继续找]]> int endPos = pos; while (endPos <= currentLen - 3) { if (fullBuf[endPos] == ']' && fullBuf[endPos+1] == ']' && fullBuf[endPos+2] == '>') { writer.write(fullBuf, pos, endPos - pos); pos = endPos + 3; break; } endPos++; } continue; } pos++; } // 剩下不够匹配最大标记长度的字符存到leftover,下次拼接处理 int remain = currentLen - pos; if (remain > 0) { if (remain >= MAX_MARKER_LEN) { writer.write(fullBuf, pos, remain - (MAX_MARKER_LEN - 1)); System.arraycopy(fullBuf, currentLen - (MAX_MARKER_LEN - 1), leftover, 0, MAX_MARKER_LEN -1); leftoverLen = MAX_MARKER_LEN -1; } else { System.arraycopy(fullBuf, pos, leftover, 0, remain); leftoverLen = remain; } } else { leftoverLen = 0; } } // 最后剩下的字符直接写入 if (leftoverLen > 0) { writer.write(leftover, 0, leftoverLen); } } }
方案2:直接提取嵌套XML内容(仅输出CDATA内的内容)
如果你的需求仅需要拿到CDATA包裹的嵌套XML,不需要外层XML结构,可以简化逻辑,匹配到<![CDATA[后开始写入内容,匹配到]]>就直接终止处理,性能更高:
private void extractNestedXml(Reader reader, String outputPath) throws IOException { final int BUFFER_SIZE = 8192; char[] buffer = new char[BUFFER_SIZE]; boolean foundStart = false; StringBuilder markerBuf = new StringBuilder(); try (BufferedWriter writer = new BufferedWriter(new FileWriter(outputPath))) { int readLen; while ((readLen = reader.read(buffer)) != -1) { for (int i = 0; i < readLen; i++) { char c = buffer[i]; if (!foundStart) { markerBuf.append(c); // 匹配到起始标记 if (markerBuf.length() >=9 && markerBuf.substring(markerBuf.length()-9).equals("<![CDATA[")) { foundStart = true; markerBuf.setLength(0); } } else { markerBuf.append(c); // 检查是否匹配结束标记 if (markerBuf.length() >=3) { if (markerBuf.substring(markerBuf.length()-3).equals("]]>")) { // 结束前写入标记之前的内容 writer.write(markerBuf.substring(0, markerBuf.length()-3)); return; } // 超过3位的部分直接写入 if (markerBuf.length() >3) { writer.write(markerBuf.charAt(0)); markerBuf.deleteCharAt(0); } } } } } // 特殊情况:没找到结束标记的话把剩余内容写入 if (foundStart && markerBuf.length() >0) { writer.write(markerBuf.toString()); } } }
注意事项
- 两个方案的内存占用仅为缓冲区大小+几字节的标记缓存,完全不需要调整JVM堆内存参数
- 全程不会修改CDATA内的嵌套XML内容,也不会额外插入换行,不会破坏嵌套XML的格式合法性
- 缓冲区大小可以根据实际情况调整,更大的缓冲区可以提升读取效率,不会带来额外的OOM风险
内容的提问来源于stack exchange,提问作者paul
相关产品推荐
相关产品推荐

