如何使用XPath提取XML空Level标签时填充空格,保证列表元素对应
XML空标签内容提取问题
场景描述
现有如下XML结构:
<IR> <name>abc</name> <Level></Level> </IR> <IR> <name>xyz</name> <Level>LVL00006</Level> </IR>
使用以下Java代码提取Level标签信息:
public class TestUseCases { public static void main(String[] args) { DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); try { factory.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true); } catch (ParserConfigurationException e1) { // TODO Auto-generated catch block e1.printStackTrace(); } Document document; try { document = factory.newDocumentBuilder() .parse(new InputSource(new StringReader(xml))); String xpathExpression = ""; Map<String, List<String>> iRValues = new LinkedHashMap<String, List<String>>(); xpathExpression = "(/IR/Level/text())"; iRValues.put("REQLVL", evaluateXPath(document, xpathExpression)); System.out.println(iRValues); } catch (SupportException e) { // TODO Auto-generated catch block e.printStackTrace(); } catch (SAXException | IOException | ParserConfigurationException e) { // TODO Auto-generated catch block e.printStackTrace(); } } private static List<String> evaluateXPath(Document document, String xpathExpression) throws SupportException { // Create XPathFactory object XPathFactory xpathFactory = XPathFactory.newInstance(); // Create XPath object XPath xpath = xpathFactory.newXPath(); List<String> values = new ArrayList<>(); try { // Create XPathExpression object XPathExpression expr = xpath.compile(xpathExpression); // Evaluate expression result on XML document NodeList nodes = (NodeList) expr.evaluate(document, XPathConstants.NODESET); System.out.println(nodes.getLength()); for (int i = 0; i < nodes.getLength(); i++) { String item = nodes.item(i).getNodeValue(); System.out.println(item); if (item.trim().isEmpty()) { values.add(SupportConstants.EMPTY_SPACE); } else { values.add(nodes.item(i).getNodeValue()); } } } catch (XPathExpressionException e) { throw new SupportException(SupportExceptionMessage.ERROR_IN_READING_RESPONSE_FROM_CM); } return values; } }
当前输出
1 LVL00006 {REQLVL=[LVL00006]}
问题与需求
第一个Level标签为空,导致提取的列表只有一个元素,无法对应到对应的name(abc/xyz)。需要实现空标签时填充空格,让最终列表变为[" ", "LVL00006"]。
解决方案
问题根源在于原XPath表达式/IR/Level/text()只会返回存在文本节点的Level标签内容,空的<Level></Level>没有text节点,因此被跳过。要解决这个问题,需要先遍历所有IR节点,再逐个处理每个IR下的Level标签:
修改后的代码
public class TestUseCases { public static void main(String[] args) { DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); try { factory.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true); } catch (ParserConfigurationException e1) { e1.printStackTrace(); } Document document; try { document = factory.newDocumentBuilder() .parse(new InputSource(new StringReader(xml))); Map<String, List<String>> iRValues = new LinkedHashMap<>(); // 修改逻辑:先获取所有IR节点再逐个处理Level List<String> levelValues = extractLevelValues(document); iRValues.put("REQLVL", levelValues); System.out.println(iRValues); } catch (SupportException e) { e.printStackTrace(); } catch (SAXException | IOException | ParserConfigurationException e) { e.printStackTrace(); } } private static List<String> extractLevelValues(Document document) throws SupportException { XPathFactory xpathFactory = XPathFactory.newInstance(); XPath xpath = xpathFactory.newXPath(); List<String> values = new ArrayList<>(); try { // 先获取所有IR节点,确保每个IR都被处理 XPathExpression irExpr = xpath.compile("/IR"); NodeList irNodes = (NodeList) irExpr.evaluate(document, XPathConstants.NODESET); // 遍历每个IR节点,处理对应的Level标签 for (int i = 0; i < irNodes.getLength(); i++) { Node irNode = irNodes.item(i); // 在当前IR节点范围内查询Level的文本 XPathExpression levelExpr = xpath.compile("Level/text()"); NodeList levelTexts = (NodeList) levelExpr.evaluate(irNode, XPathConstants.NODESET); String levelValue; // 无文本节点或文本为空时填充指定空格 if (levelTexts.getLength() == 0 || levelTexts.item(0).getNodeValue().trim().isEmpty()) { levelValue = SupportConstants.EMPTY_SPACE; } else { levelValue = levelTexts.item(0).getNodeValue(); } values.add(levelValue); } } catch (XPathExpressionException e) { throw new SupportException(SupportExceptionMessage.ERROR_IN_READING_RESPONSE_FROM_CM); } return values; } }
说明
- 先通过
/IR获取所有IR节点,确保每个IR都会被处理,不会遗漏空Level的情况; - 对每个IR节点,在其范围内查询Level的文本内容,避免跨节点匹配;
- 如果Level没有文本节点(空标签)或文本内容为空,则填充
SupportConstants.EMPTY_SPACE,否则使用原文本; - 最终返回的列表严格遵循XML中IR的顺序,每个IR对应一个元素,实现
[" ", "LVL00006"]的需求。
内容的提问来源于stack exchange,提问作者Ajitesh Shukla
相关产品推荐
相关产品推荐

