You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JAXP解析合法XML出错,Java XPath查询遇阻

JDK 1.8 DocumentBuilder解析含单个]的CDATA XML报错的解决方案

我刚帮你排查了这个问题——你说的完全没错,根据XML规范,CDATA段里单个]出现在]]>之前是完全合法的,问题出在JDK 1.8自带的XML解析器(Xerces 2.11.0版本)的一个已知bug上,这个版本的解析器会错误地把CDATA内的单个]和结尾的]]>混淆,导致抛出解析异常。

先给你确认下问题重现的场景:

最小化错误XML示例

<root>
  <data><![CDATA[Some text with a single ] before the end]]></data>
</root>

常规JAXP解析代码(会报错)

import javax.xml.parsers.DocumentBuilder;
import javax.xml.parsers.DocumentBuilderFactory;
import org.w3c.dom.Document;
import java.io.ByteArrayInputStream;

public class XmlParseTest {
    public static void main(String[] args) throws Exception {
        String xml = "<root><data><![CDATA[Some text with a single ] before the end]]></data></root>";
        DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
        DocumentBuilder builder = factory.newDocumentBuilder();
        // 执行到这一步会抛出SAXParseException
        Document doc = builder.parse(new ByteArrayInputStream(xml.getBytes()));
    }
}

解决办法

1. 升级XML解析器版本

JDK 1.8自带的Xerces版本过低,你可以手动引入更新版的Xerces库(比如2.12.2及以上),以Maven项目为例,添加依赖:

<dependency>
    <groupId>xerces</groupId>
    <artifactId>xercesImpl</artifactId>
    <version>2.12.2</version>
</dependency>

然后在代码中指定使用这个解析器:

// 在创建DocumentBuilderFactory前设置系统属性
System.setProperty("javax.xml.parsers.DocumentBuilderFactory", 
                   "org.apache.xerces.jaxp.DocumentBuilderFactoryImpl");

DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
// 后续解析代码不变

2. 使用其他XML解析库

如果不想依赖Xerces,也可以改用DOM4J或者JDOM2这类第三方库,它们对CDATA的处理更规范。比如DOM4J的示例代码:

import org.dom4j.Document;
import org.dom4j.io.SAXReader;
import java.io.ByteArrayInputStream;

public class Dom4jParseDemo {
    public static void main(String[] args) throws Exception {
        String xml = "<root><data><![CDATA[Some text with a single ] before the end]]></data></root>";
        SAXReader reader = new SAXReader();
        Document doc = reader.read(new ByteArrayInputStream(xml.getBytes()));
        // 成功解析,后续可以正常执行XPath查询
        String content = doc.selectSingleNode("//data").getText();
        System.out.println(content);
    }
}

3. 升级JDK版本

如果条件允许,直接升级到JDK 11及以上版本,这些版本自带的解析器已经修复了这个CDATA处理的bug,无需额外配置即可正常解析。

补充说明

你提到用Saxon API也遇到了相同错误,大概率是因为你没有正确配置Saxon作为解析器。如果要使用Saxon,需要确保引入了Saxon的依赖,并指定使用它的DocumentBuilderFactory:

DocumentBuilderFactory factory = new net.sf.saxon.jaxp.DocumentBuilderFactoryImpl();

这样Saxon的解析器会正确处理CDATA段中的单个]字符。

内容的提问来源于stack exchange,提问作者Robert Jack Will

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:02:17