Java XPath无法定位XML Schema中simpleType元素问题排查
问题:XPath无法定位XML Schema中的元素及相关异常
我现在碰到一个XPath定位XML元素的问题,没法从给定的XML Schema里找到第一个<simpleType>元素。而且测试XPath表达式的时候还出现了不符合预期的结果,麻烦帮忙排查一下。
给定的XML Schema
<?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema" elementFormDefault="qualified"> <xs:simpleType name="IdType"> <xs:annotation> <xs:documentation>Уникальный идентификатор (ключ) объекта в АИС контрагента</xs:documentation> </xs:annotation> <xs:restriction base="xs:token"> <xs:minLength value="1"/> <xs:maxLength value="60"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="DateType"> <xs:annotation> <xs:documentation>Дата</xs:documentation> </xs:annotation> <xs:restriction base="xs:date"> <xs:minInclusive value="1900-01-01"/> <xs:maxExclusive value="2100-01-01"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="DateTimeType"> <xs:annotation> <xs:documentation>Дата и время</xs:documentation> </xs:annotation> <xs:restriction base="xs:dateTime"> <xs:minInclusive value="1900-01-01T00:00:00"/> <xs:maxExclusive value="2100-01-01T00:00:00"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="String_1_1000"> <xs:annotation> <xs:documentation>Строка длиной от 1 до 1000 символов</xs:documentation> </xs:annotation> <xs:restriction base="xs:string"> <xs:minLength value="1"/> <xs:maxLength value="1000"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="DocumentNumberType"> <xs:annotation> <xs:documentation>Номер документа</xs:documentation> </xs:annotation> <xs:restriction base="xs:string"> <xs:minLength value="1"/> <xs:maxLength value="25"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="OrganizationCodeType"> <xs:annotation> <xs:documentation>Код организации</xs:documentation> </xs:annotation> <xs:restriction base="xs:string"> <xs:minLength value="1"/> <xs:maxLength value="32"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="DepartmentCodeType"> <xs:annotation> <xs:documentation>Код подразделения</xs:documentation> </xs:annotation> <xs:restriction base="xs:string"> <xs:minLength value="1"/> <xs:maxLength value="32"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="CaseNumberType"> <xs:annotation> <xs:documentation>Номер дела</xs:documentation> </xs:annotation> <xs:restriction base="xs:string"> <xs:minLength value="1"/> <xs:maxLength value="25"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="DocumentDataType"> <xs:annotation> <xs:documentation>Тип документа электронного документооборота</xs:documentation> </xs:annotation> <xs:restriction base="xs:string"> <xs:minLength value="1"/> <xs:maxLength value="32"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="StateIdType"> <xs:annotation> <xs:documentation>Состояние документа</xs:documentation> </xs:annotation> <xs:restriction base="xs:string"> <xs:enumeration value="90"/> <xs:enumeration value="91"/> <xs:enumeration value="92"/> <xs:enumeration value="93"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="AttachmentFilenameType"> <xs:annotation> <xs:documentation>Имя файла вложения</xs:documentation> </xs:annotation> <xs:restriction base="xs:ID"> <xs:maxLength value="1024"/> </xs:restriction> </xs:simpleType> </xs:schema>
我的Java解析代码
public Node getFirstByExpression(final String expr) { DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); factory.setNamespaceAware(false); try { DocumentBuilder builder = factory.newDocumentBuilder(); document = builder.parse(source); document.getDocumentElement().normalize(); XPathFactory factory = XPathFactory.newInstance(); XPath xPath = factory.newXPath(); NodeList nodeList = null; XPathExpression expression = xPath.compile(expr); nodeList = (NodeList) expression.evaluate(document, XPathConstants.NODESET); return nodeList.getLength() == 0 ? null : nodeList.item(0); } catch (Exception ex) { log.error(ex); return null; } }
测试结果
assert parser.getFirstByExpression("schema") != null;// 结果为trueassert parser.getFirstByExpression("/schema") != null;// 结果为trueassert parser.getFirstByExpression("//schema") != null;// 结果为false,预期应为true
当传入参数"simpleType"时,nodeList长度为0,方法返回null,完全找不到目标元素。
问题排查与解决方案
核心问题出在XML命名空间的处理上:你的XML Schema里所有元素都属于http://www.w3.org/2001/XMLSchema命名空间,使用了xs前缀,但你的代码关闭了命名空间感知,且XPath表达式没有考虑命名空间。
1. 为什么当前代码会出现异常结果?
当factory.setNamespaceAware(false);时,DOM解析器不会解析XML的命名空间,元素的节点名会被当作包含前缀的完整字符串(比如xs:schema、xs:simpleType)。此时XPath表达式中的schema或simpleType无法匹配到带前缀的节点,也就出现了找不到元素、断言不符合预期的情况。
2. 修复步骤:
步骤1:开启命名空间感知
把factory.setNamespaceAware(false);修改为:
factory.setNamespaceAware(true);
步骤2:配置XPath的命名空间上下文
为XPath绑定xs前缀到对应的命名空间URI,这样表达式就能正确识别带前缀的元素:
xPath.setNamespaceContext(new NamespaceContext() { @Override public String getNamespaceURI(String prefix) { if ("xs".equals(prefix)) { return "http://www.w3.org/2001/XMLSchema"; } return XMLConstants.NULL_NS_URI; } @Override public String getPrefix(String namespaceURI) { return null; // 不需要实现反向查找逻辑 } @Override public Iterator<String> getPrefixes(String namespaceURI) { return null; // 不需要实现反向查找逻辑 } });
步骤3:修改XPath表达式为带前缀的形式
- 要找第一个
<simpleType>,使用表达式:"//xs:simpleType" - 测试
schema的表达式改为:"//xs:schema",此时断言会返回true - 其他元素的查找也要带上
xs:前缀,比如"xs:schema/xs:simpleType"
3. 修复后的完整代码示例
import javax.xml.namespace.NamespaceContext; import javax.xml.xpath.XPath; import javax.xml.xpath.XPathExpression; import javax.xml.xpath.XPathFactory; import org.w3c.dom.Node; import org.w3c.dom.NodeList; import javax.xml.XMLConstants; public Node getFirstByExpression(final String expr) { DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); factory.setNamespaceAware(true); // 开启命名空间感知 try { DocumentBuilder builder = factory.newDocumentBuilder(); document = builder.parse(source); document.getDocumentElement().normalize(); XPathFactory xPathFactory = XPathFactory.newInstance(); XPath xPath = xPathFactory.newXPath(); // 配置命名空间上下文 xPath.setNamespaceContext(new NamespaceContext() { @Override public String getNamespaceURI(String prefix) { if ("xs".equals(prefix)) { return "http://www.w3.org/2001/XMLSchema"; } return XMLConstants.NULL_NS_URI; } @Override public String getPrefix(String namespaceURI) { return null; } @Override public Iterator<String> getPrefixes(String namespaceURI) { return null; } }); XPathExpression expression = xPath.compile(expr); NodeList nodeList = (NodeList) expression.evaluate(document, XPathConstants.NODESET); return nodeList.getLength() == 0 ? null : nodeList.item(0); } catch (Exception ex) { log.error(ex); return null; } }
4. 测试验证
现在使用以下表达式就能得到预期结果:
parser.getFirstByExpression("//xs:schema")// 返回非null,符合预期parser.getFirstByExpression("//xs:simpleType")// 返回第一个<xs:simpleType>元素,不再是null
内容的提问来源于stack exchange,提问作者nllsdfx
相关产品推荐
相关产品推荐

