如何使用Java XPath实现XML文件全文件关键词检索并判断匹配结果
基于XPath的XML全量搜索实现方案
需求背景
我正在开发一款应用,该应用会接收用户输入的关键词,随后通过XPath扫描XML文件,根据关键词是否在文件中存在返回true或false结果。
最初实现的问题代码
public XPath() throws IOException, SAXException, ParserConfigurationException, XPathExpressionException { FileInputStream fileIS = new FileInputStream("text.xml"); DocumentBuilderFactory builderFactory = DocumentBuilderFactory.newInstance(); DocumentBuilder builder = builderFactory.newDocumentBuilder(); Document xmlDocument = builder.parse(fileIS); XPathFactory xPathfactory = XPathFactory.newInstance(); javax.xml.xpath.XPath xPath = xPathfactory.newXPath(); XPathExpression expr = xPath.compile("//text()[contains(.,'java')]"); System.out.println(expr.evaluate(xmlDocument, XPathConstants.NODESET)); }
测试用XML文件内容
<?xml version="1.0"?> <Tutorials> <Tutorial tutId="01" type="java"> <title>Guava</title> <description>Introduction to Guava</description> <date>04/04/2016</date> <author>GuavaAuthor</author> </Tutorial> <Tutorial tutId="02" type="java"> <title>XML</title> <description>Introduction to XPath</description> <date>04/05/2016</date> <author>XMLAuthor</author> </Tutorial> </Tutorials>
问题原因与解决方案
问题分析
- 原XPath表达式
//text()[contains(.,'java')]仅扫描XML的文本节点,但测试文件中java关键词仅存在于Tutorial节点的type属性值中,因此永远无法匹配到结果 - 原代码直接打印XPath返回的节点集对象,缺少遍历节点取内容的逻辑,无法正常展示匹配结果
修复方案
- 可根据搜索需求调整XPath表达式:
- 仅搜索文本节点:保持原表达式,更换存在于文本内容中的测试关键词即可正常匹配
- 全量搜索(同时匹配文本内容和所有属性值):可修改表达式为
//*[contains(text(),'${替换为实际搜索关键词}') or @*[contains(.,'${替换为实际搜索关键词}')]]
- 新增节点集遍历输出逻辑,若仅需要判断关键词是否存在,可直接返回节点集长度是否大于0的布尔值,代码如下:
Object result = expr.evaluate(xmlDocument, XPathConstants.NODESET); NodeList nodes = (NodeList) result; for (int i = 0; i < nodes.getLength(); i++) { System.out.println(nodes.item(i).getNodeValue()); } // 判断是否存在匹配的代码:return nodes.getLength() > 0;
内容的提问来源于stack exchange,提问作者Adrian
相关产品推荐
相关产品推荐

