Xerces2-J与Oxygen Editor的XML验证错误定位精度差异及实现方式咨询
你遇到的这个问题其实挺典型的——同样基于Xerces引擎,Oxygen能精准定位到错误文本的位置,而原生Xerces-J的默认输出却指向父元素的结束标签。我来给你拆解背后的原因和可行的解决方案:
先回顾下你的测试场景
测试用XSD:
<?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"> <xs:element name="shiporder"> <xs:complexType> <xs:sequence> <xs:element name="orderperson" type="xs:string"/> <xs:element name="shipto"> <xs:complexType> <xs:sequence> <xs:element name="name" type="xs:string"/> <xs:element name="address" type="xs:string"/> <xs:element name="city" type="xs:string"/> <xs:element name="country" type="xs:string"/> </xs:sequence> </xs:complexType> </xs:element> <xs:element name="item" maxOccurs="unbounded"> <xs:complexType> <xs:sequence> <xs:element name="title" type="xs:string"/> <xs:element name="note" type="xs:string" minOccurs="0"/> <xs:element name="quantity" type="xs:positiveInteger"/> <xs:element name="price" type="xs:decimal"/> </xs:sequence> </xs:complexType> </xs:element> </xs:sequence> <xs:attribute name="orderid" type="xs:string" use="required"/> </xs:complexType> </xs:element> </xs:schema>
测试用XML:
<?xml version="1.0" encoding="UTF-8"?> <shiporder xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" orderid="889923" xsi:noNamespaceSchemaLocation="shiporder.xsd"> <orderperson>John Smith</orderperson> <shipto>--- Bogus Text --- <name>Ola Nordmann</name> <address>Langgt 23</address> <city>4000 Stavanger</city> <country>Norway</country> </shipto> <item> <title>Empire Burlesque</title> <note>Special Edition</note> <quantity>1</quantity> <price>10.90</price> </item> <item> <title>Hide your heart</title> <quantity>1</quantity> <price>9.90</price> </item> </shiporder>
错误对比:
- 原生Xerces-J 2.12.2的输出:
[Error] doc.xml:9:12: cvc-complex-type.2.3: Element 'shipto' cannot have character [children], because the type's content type is element-only.
- Oxygen Editor的输出:
Start location: line: 4, column: 11 End location: line: 4, column: 29
为什么会有这个差异?
Oxygen确实是对Xerces做了定制化扩展:
- 原生Xerces在验证element-only类型的元素时,当遇到非法的字符节点,默认会把错误关联到父元素的结束标签位置——这是因为它在验证时,是等父元素解析完成后才抛出整体的内容模型错误。
- 而Oxygen在Xerces的基础上,额外实现了更精细的位置追踪逻辑:它会实时记录每个字符节点的起始/结束位置,当触发内容模型错误时,直接定位到非法文本的具体位置,而不是父元素的结束处。
原生Xerces-J能不能实现精准定位?
答案是可以!你需要做两个调整:
1. 开启Xerces的完整验证特性
在创建Validator时,启用schema-full-checking特性,这个特性会让Xerces做更细致的验证检查,包括更精准的位置追踪:
SchemaFactory factory = SchemaFactory.newInstance(XMLConstants.W3C_XML_SCHEMA_NS_URI); Schema schema = factory.newSchema(new File("shiporder.xsd")); Validator validator = schema.newValidator(); // 开启完整模式检查 validator.setFeature("http://apache.org/xml/features/validation/schema-full-checking", true);
2. 自定义ErrorHandler提取详细位置
原生Xerces的SAXParseException其实包含了起始和结束位置的信息,但默认的错误输出只打印了结束位置。你可以自定义ErrorHandler来获取完整的位置数据:
validator.setErrorHandler(new ErrorHandler() { @Override public void warning(SAXParseException exception) throws SAXException { printDetailedError(exception); } @Override public void error(SAXParseException exception) throws SAXException { printDetailedError(exception); } @Override public void fatalError(SAXParseException exception) throws SAXException { printDetailedError(exception); } private void printDetailedError(SAXParseException e) { System.out.printf("[Error] %s:%d:%d to %d:%d: %s%n", e.getSystemId(), e.getLineNumber(), e.getColumnNumber(), e.getEndLineNumber(), // 关键:获取结束行号 e.getEndColumnNumber(), // 关键:获取结束列号 e.getMessage()); } });
通过getEndLineNumber()和getEndColumnNumber()方法,你就能拿到错误的完整位置范围,和Oxygen的输出一致。
总结一下:Oxygen是定制了Xerces来优化错误体验,但原生Xerces-J本身就支持获取精准的错误位置,只要你开启对应的特性并自定义ErrorHandler来提取详细信息就行。
内容来源于stack exchange

