如何用Python的lxml模块结合XPath验证title节点是否为section1的后代?
解决XPath验证节点祖先的错误问题
错误原因
你编写的XPath表达式. [ancestor::/library/section1]违反语法规则:ancestor::轴后面需要直接跟节点名称(节点测试),不能带路径斜杠,这才导致了XPathEvalError: Invalid expression报错。
正确的验证方法
以下两种XPath写法都可以实现需求:
方法一:精准匹配祖先层级
检查当前title节点是否存在祖先section1,且该section1的父节点是根节点library:
titles = root.xpath('//title') for title in titles: if title.xpath('.[ancestor::section1[parent::library]]'): do_sth() # 执行你的逻辑
方法二:简洁路径写法
直接验证当前节点是否在library/section1的后代路径中:
titles = root.xpath('//title') for title in titles: if title.xpath('.[ancestor::library/section1]'): do_sth() # 执行你的逻辑
完整示例代码
from lxml import etree # 解析目标XML xml_content = """ <library> <section1> <book> <title>Harry Potter</title> <author>J.K. Rowling</author> </book> </section1> <section2> <book> <title>Sapiens</title> <author>Yuval Noah Harari</author> </book> </section2> </library> """ root = etree.fromstring(xml_content) # 获取所有title节点 titles = root.xpath('//title') # 遍历验证并处理 for title in titles: if title.xpath('.[ancestor::library/section1]'): print(f"匹配到section1下的标题:{title.text}") # do_sth() # 替换为你的业务逻辑
内容的提问来源于stack exchange,提问作者zsfzu0
相关产品推荐
相关产品推荐

