You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用XPath提取XML空Level标签时填充空格,保证列表元素对应

XML空标签内容提取问题

场景描述

现有如下XML结构:

<IR>
        <name>abc</name>
        <Level></Level>
      
</IR>
<IR>
        <name>xyz</name> 
        <Level>LVL00006</Level>     
</IR>

使用以下Java代码提取Level标签信息:

public class TestUseCases {
 public static void main(String[] args) {
    DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
    try {
        factory.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
    } catch (ParserConfigurationException e1) {
        // TODO Auto-generated catch block
        e1.printStackTrace();
    }
    Document document;
    try {
        document = factory.newDocumentBuilder()
                .parse(new InputSource(new StringReader(xml)));
    
    String xpathExpression = "";
    Map<String, List<String>> iRValues = new LinkedHashMap<String, List<String>>();
    xpathExpression = "(/IR/Level/text())";
    
        iRValues.put("REQLVL", evaluateXPath(document, xpathExpression));
        
        System.out.println(iRValues);
    } catch (SupportException e) {
        // TODO Auto-generated catch block
        e.printStackTrace();
    } catch (SAXException | IOException | ParserConfigurationException e) {
        // TODO Auto-generated catch block
        e.printStackTrace();
    }
  }

  private static List<String> evaluateXPath(Document document, String xpathExpression) throws SupportException {
    // Create XPathFactory object
    XPathFactory xpathFactory = XPathFactory.newInstance();

    // Create XPath object
    XPath xpath = xpathFactory.newXPath();

    List<String> values = new ArrayList<>();
    try {
        // Create XPathExpression object
        XPathExpression expr = xpath.compile(xpathExpression);
        // Evaluate expression result on XML document
        NodeList nodes = (NodeList) expr.evaluate(document, XPathConstants.NODESET);
        System.out.println(nodes.getLength());
        for (int i = 0; i < nodes.getLength(); i++) {
            String item = nodes.item(i).getNodeValue();
            System.out.println(item);
            if (item.trim().isEmpty()) {
                values.add(SupportConstants.EMPTY_SPACE);
            } else {
                values.add(nodes.item(i).getNodeValue());
            }
        }

    } catch (XPathExpressionException e) {
        throw new SupportException(SupportExceptionMessage.ERROR_IN_READING_RESPONSE_FROM_CM);
    }
    return values;
  }
}

当前输出

1
LVL00006
{REQLVL=[LVL00006]}

问题与需求

第一个Level标签为空,导致提取的列表只有一个元素,无法对应到对应的name(abc/xyz)。需要实现空标签时填充空格,让最终列表变为[" ", "LVL00006"]。


解决方案

问题根源在于原XPath表达式/IR/Level/text()只会返回存在文本节点的Level标签内容,空的<Level></Level>没有text节点,因此被跳过。要解决这个问题,需要先遍历所有IR节点,再逐个处理每个IR下的Level标签:

修改后的代码

public class TestUseCases {
 public static void main(String[] args) {
    DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
    try {
        factory.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
    } catch (ParserConfigurationException e1) {
        e1.printStackTrace();
    }
    Document document;
    try {
        document = factory.newDocumentBuilder()
                .parse(new InputSource(new StringReader(xml)));
    
    Map<String, List<String>> iRValues = new LinkedHashMap<>();
    // 修改逻辑:先获取所有IR节点再逐个处理Level
    List<String> levelValues = extractLevelValues(document);
    iRValues.put("REQLVL", levelValues);
        
    System.out.println(iRValues);
    } catch (SupportException e) {
        e.printStackTrace();
    } catch (SAXException | IOException | ParserConfigurationException e) {
        e.printStackTrace();
    }
  }

  private static List<String> extractLevelValues(Document document) throws SupportException {
    XPathFactory xpathFactory = XPathFactory.newInstance();
    XPath xpath = xpathFactory.newXPath();
    List<String> values = new ArrayList<>();
    
    try {
        // 先获取所有IR节点,确保每个IR都被处理
        XPathExpression irExpr = xpath.compile("/IR");
        NodeList irNodes = (NodeList) irExpr.evaluate(document, XPathConstants.NODESET);
        
        // 遍历每个IR节点,处理对应的Level标签
        for (int i = 0; i < irNodes.getLength(); i++) {
            Node irNode = irNodes.item(i);
            // 在当前IR节点范围内查询Level的文本
            XPathExpression levelExpr = xpath.compile("Level/text()");
            NodeList levelTexts = (NodeList) levelExpr.evaluate(irNode, XPathConstants.NODESET);
            
            String levelValue;
            // 无文本节点或文本为空时填充指定空格
            if (levelTexts.getLength() == 0 || levelTexts.item(0).getNodeValue().trim().isEmpty()) {
                levelValue = SupportConstants.EMPTY_SPACE;
            } else {
                levelValue = levelTexts.item(0).getNodeValue();
            }
            values.add(levelValue);
        }

    } catch (XPathExpressionException e) {
        throw new SupportException(SupportExceptionMessage.ERROR_IN_READING_RESPONSE_FROM_CM);
    }
    return values;
  }
}

说明

  1. 先通过/IR获取所有IR节点,确保每个IR都会被处理,不会遗漏空Level的情况;
  2. 对每个IR节点,在其范围内查询Level的文本内容,避免跨节点匹配;
  3. 如果Level没有文本节点(空标签)或文本内容为空,则填充SupportConstants.EMPTY_SPACE,否则使用原文本;
  4. 最终返回的列表严格遵循XML中IR的顺序,每个IR对应一个元素,实现[" ", "LVL00006"]的需求。

内容的提问来源于stack exchange,提问作者Ajitesh Shukla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 22:16:39