如何用XPath递归解析任意深度XML为List<Map<String,String>>(Java)
在Java中能否将任意深度的XML递归转换为List<Map<String,String>>?
我目前需要把任意层级的XML数据转换成List<Map<String,String>>格式,其中Map的键是XML节点的层级路径(比如address.street),值是节点的文本内容。我自己写的代码只能处理2层深度,想请教有没有办法支持任意深度的解析?
我的XML数据示例:
<?xml version="1.0" encoding="UTF-8"?> <p:PersonalDetails> <Node_1> <Node_1_1> <name>name 1</name> <address> <street>17</street> <town>1507487</town> </address> <details> <detail_1>detaile item 1</detail_1> <detail_2> <detail_2_1>detail item 2_1</detail_2_1> <detail_2_2>detail item 2_1</detail_2_2> </detail_2> </details> </Node_1_1> <Node_1_2> <name>name 1</name> <address> <street>17</street> <town>1507487</town> </address> <details> <detail_1> <detail_1_1> <detail_1_1_1>detail item 2_1_1</detail_1_1_1> </detail_1_1> <detail_1_2>detail item 2_1</detail_1_2> </detail_1> <detail_2> <detail_2_1> <detail_2_1_1> <detail_2_1_1_1>detail item 2_1_1_1</detail_2_1_1_1> </detail_2_1_1> </detail_2_1> </detail_2> </details> </Node_1_2> </Node_1> </p:PersonalDetails>
我现有的代码(仅支持2层深度):
public static void testXpath(String filePath, String expr,String childSubNodeName) throws ParserConfigurationException, XPathExpressionException, IOException, SAXException { DocumentBuilderFactory builderFactory = DocumentBuilderFactory.newInstance(); DocumentBuilder builder = builderFactory.newDocumentBuilder(); Document xmlDocument = builder.parse(filePath); xmlDocument.getDocumentElement().normalize(); XPath xPath = XPathFactory.newInstance().newXPath(); NodeList nodeList = (NodeList) xPath.compile("//"+expr).evaluate(xmlDocument, XPathConstants.NODESET); List<Map<String,String>> listMap = new LinkedList<>(); for(int i=0;i<nodeList.getLength();i++){ NodeList childNode = (NodeList) nodeList.item(i); Map<String,String> map = new HashMap<>(); for(int j=0;j<childNode.getLength();j++){ if(!childNode.item(j).getTextContent().equals("\n")){ if(childNode.item(j).getNodeName().contains(childSubNodeName)) { extractSubNode(childNode.item(j), map); } else map.put(childNode.item(j).getNodeName(), childNode.item(j).getTextContent()); } } listMap.add(map); } System.out.println(listMap); System.out.println("-------------------------"); } private static void extractSubNode(Node item, Map<String, String> map) { NodeList subNode = item.getChildNodes(); for(int j=0;j<subNode.getLength();j++){ if(!subNode.item(j).getTextContent().equals("\n")){ map.put(item.getNodeName()+"."+subNode.item(j).getNodeName(),subNode.item(j).getTextContent()); } } }
我期望的输出格式:
[{name=name 1, address.street=17, address.town=1507487, details.detail_1=detaile item 1, details.detail_2.detail_2_1=detail item 2_1, details.detail_2.detail_2_2=detail item 2_1}, ...]
解决方案
完全可以实现任意深度的XML解析,核心是把节点解析逻辑改成递归方法,每次处理子节点时将当前节点的路径作为前缀传递下去,直到遇到最终的文本节点为止。
修改后的完整代码:
import org.w3c.dom.Document; import org.w3c.dom.Node; import org.w3c.dom.NodeList; import javax.xml.parsers.DocumentBuilder; import javax.xml.parsers.DocumentBuilderFactory; import javax.xml.xpath.XPath; import javax.xml.xpath.XPathConstants; import javax.xml.xpath.XPathFactory; import java.util.HashMap; import java.util.LinkedList; import java.util.List; import java.util.Map; public class XmlToMapConverter { public static void testXpath(String filePath, String expr) throws Exception { DocumentBuilderFactory builderFactory = DocumentBuilderFactory.newInstance(); // 开启命名空间支持,适配带命名空间的XML builderFactory.setNamespaceAware(true); DocumentBuilder builder = builderFactory.newDocumentBuilder(); Document xmlDocument = builder.parse(filePath); xmlDocument.getDocumentElement().normalize(); XPath xPath = XPathFactory.newInstance().newXPath(); NodeList nodeList = (NodeList) xPath.compile("//" + expr).evaluate(xmlDocument, XPathConstants.NODESET); List<Map<String, String>> listMap = new LinkedList<>(); for (int i = 0; i < nodeList.getLength(); i++) { Node currentNode = nodeList.item(i); Map<String, String> map = new HashMap<>(); // 递归解析当前节点的所有子节点,初始前缀为空 extractRecursiveNodes(currentNode, "", map); listMap.add(map); } System.out.println(listMap); } /** * 递归解析XML节点,拼接层级路径作为Map的键 * @param node 当前处理的XML节点 * @param prefix 父节点的路径前缀 * @param map 存储结果的Map */ private static void extractRecursiveNodes(Node node, String prefix, Map<String, String> map) { NodeList childNodes = node.getChildNodes(); for (int j = 0; j < childNodes.getLength(); j++) { Node child = childNodes.item(j); // 跳过空白文本节点和注释节点 if (child.getNodeType() == Node.TEXT_NODE || child.getNodeType() == Node.COMMENT_NODE) { String text = child.getTextContent().trim(); if (!text.isEmpty() && !prefix.isEmpty()) { map.put(prefix, text); } continue; } // 拼接新的路径前缀 String newPrefix = prefix.isEmpty() ? child.getNodeName() : prefix + "." + child.getNodeName(); // 递归处理下一层子节点 extractRecursiveNodes(child, newPrefix, map); } } public static void main(String[] args) throws Exception { // 示例调用:解析Node_1下的所有子节点(Node_1_1和Node_1_2) testXpath("your-xml-file-path.xml", "Node_1/*"); } }
关键改进点
- 递归遍历:
extractRecursiveNodes方法会逐层遍历所有子节点,支持任意深度的XML结构 - 路径拼接:每次递归传递当前节点的路径前缀,最终生成
address.street这类层级化键名 - 节点过滤:自动跳过空白文本和注释节点,避免生成无效键值对
- 命名空间兼容:开启
setNamespaceAware(true),适配带命名空间的XML文档
运行后即可得到你期望的任意深度List<Map<String,String>>格式结果。
内容的提问来源于stack exchange,提问作者Piya
相关产品推荐
相关产品推荐

