如何优雅解析ZipInputStream中的多个XML条目?
我实现了一个导出功能,将两个XML打包成ZIP供下载。单元测试时需要验证导出结果,用ZipInputStream读取ZIP并解析XML。最初用IOUtils.toString()把当前条目转成字符串再解析是正常的,但尝试直接用InputSource(new InputStreamReader(zis))时,第一个条目正常,第二个条目抛出异常——原因是XML解析器会读取到流的末尾,甚至关闭底层的ZipInputStream,导致后续条目无法读取。
解决方案1:限制解析流的读取长度
利用ZipEntry.getSize()获取当前条目的大小,把ZipInputStream包装成一个仅读取指定长度的流,避免解析器读超。如果使用Apache Commons IO,可以用BoundedInputStream:
NodeList getNodesByName(ZipInputStream zis, String nodeName) throws IOException, ParserConfigurationException, SAXException { ZipEntry currentEntry = zis.getCurrentEntry(); // 包装成限制长度的流,禁止关闭底层ZipInputStream BoundedInputStream boundedStream = new BoundedInputStream(zis, currentEntry.getSize()); boundedStream.setPropagateClose(false); InputSource is = new InputSource(new InputStreamReader(boundedStream)); Document doc = DocumentBuilderFactory.newInstance().newDocumentBuilder().parse(is); return doc.getElementsByTagName(nodeName); }
如果不想依赖第三方库,可以先把当前条目内容读入字节数组,再用内存流解析:
NodeList getNodesByName(ZipInputStream zis, String nodeName) throws IOException, ParserConfigurationException, SAXException { // 读取当前条目全部内容到字节数组 byte[] entryContent = IOUtils.toByteArray(zis); try (ByteArrayInputStream byteStream = new ByteArrayInputStream(entryContent)) { InputSource is = new InputSource(new InputStreamReader(byteStream)); Document doc = DocumentBuilderFactory.newInstance().newDocumentBuilder().parse(is); return doc.getElementsByTagName(nodeName); } }
这种方式和最初的实现逻辑类似,但直接操作字节数组更高效,且不会影响底层ZipInputStream的状态。
解决方案2:改用ZipFile处理(更推荐)
ZipFile可以单独获取每个条目的输入流,每个条目流相互独立,解析器可以安全读取到流末尾而不影响其他条目。需要先把导出的ZIP流写入临时文件:
// 把导出的ZIP流写入临时文件 File tempZip = File.createTempFile("export_test", ".zip"); try (OutputStream os = new FileOutputStream(tempZip)) { IOUtils.copy(testee.openStream(), os); } // 用ZipFile读取每个条目 try (ZipFile zipFile = new ZipFile(tempZip)) { // 处理第一个XML ZipEntry firstEntry = zipFile.getEntry("first.xml"); // 替换为实际文件名 assertNotNull(firstEntry); try (InputStream entryStream = zipFile.getInputStream(firstEntry)) { NodeList nodesOne = getNodesByName(entryStream, "nodeOne"); assertEquals(EXPECTED_NODE_ONE, nodesOne.getLength()); } // 处理第二个XML ZipEntry secondEntry = zipFile.getEntry("second.xml"); // 替换为实际文件名 assertNotNull(secondEntry); try (InputStream entryStream = zipFile.getInputStream(secondEntry)) { NodeList nodesTwo = getNodesByName(entryStream, "nodeTwo"); assertEquals(EXPECTED_NODE_TWO, nodesTwo.getLength()); } } finally { // 清理临时文件 tempZip.delete(); } // 此时getNodesByName可以简化为通用的流解析方法 NodeList getNodesByName(InputStream is, String nodeName) throws ParserConfigurationException, IOException, SAXException { InputSource is = new InputSource(new InputStreamReader(is)); Document doc = DocumentBuilderFactory.newInstance().newDocumentBuilder().parse(is); return doc.getElementsByTagName(nodeName); }
问题根源
DocumentBuilder.parse()方法会读取整个输入流直到EOF,而ZipInputStream的EOF是整个ZIP文件的末尾,不是当前条目的末尾。解析第一个条目时,解析器会把后续条目甚至ZIP文件的结束标识都读入,导致ZipInputStream的指针直接跳到文件末尾,后续getNextEntry()无法找到有效条目。部分解析器实现还会关闭输入流,直接导致ZipInputStream失效。
内容的提问来源于stack exchange,提问作者Thomas

