Python中如何传递XPath表达式及通过标签名生成XPath?
在Python中用XPath解析本地XML并生成XPath表达式
一、用XPath提取本地XML数据
Python里处理XPath最常用的是lxml库(对XPath语法支持更完整,比标准库的xml.etree.ElementTree功能更强),操作步骤如下:
- 安装lxml依赖
pip install lxml
- 加载本地XML并执行XPath查询
假设你的XML文件为data.xml,内容示例:
<library> <book category="tech"> <title>Python核心编程</title> <author>王五</author> </book> <book category="literature"> <title>百年孤独</title> <author>加西亚·马尔克斯</author> </book> </library>
代码示例:
from lxml import etree # 加载本地XML文件 tree = etree.parse("data.xml") # 提取所有book节点下的title文本 titles = tree.xpath("//book/title/text()") for title in titles: print(title) # 提取category为tech的book的author tech_authors = tree.xpath("//book[@category='tech']/author/text()") print(tech_authors)
如果想用标准库xml.etree.ElementTree(仅支持XPath 1.0子集),代码如下:
import xml.etree.ElementTree as ET tree = ET.parse("data.xml") root = tree.getroot() # 查找所有title元素 titles = root.findall(".//title") for title in titles: print(title.text)
二、通过标签名生成XPath表达式
可以写一个轻量化函数,根据传入的标签名和可选参数生成不同场景的XPath:
基础版:支持父节点、属性匹配
def generate_xpath(tag_name, parent_tag=None, attr_match=None): # 默认匹配所有该标签 xpath = f"//{tag_name}" # 限定父节点路径 if parent_tag: xpath = f"//{parent_tag}/{tag_name}" # 添加属性匹配条件 if attr_match: # attr_match为字典,比如{"category": "tech"} attr_str = " and ".join([f"@{k}='{v}'" for k, v in attr_match.items()]) xpath = f"{xpath}[{attr_str}]" return xpath # 示例调用 print(generate_xpath("title")) # 输出://title print(generate_xpath("title", parent_tag="book")) # 输出://book/title print(generate_xpath("book", attr_match={"category": "literature"})) # 输出://book[@category='literature']
进阶版:支持文本匹配、位置筛选
def generate_xpath(tag_name, parent_tag=None, text_contains=None, position=None): xpath = f"//{tag_name}" if parent_tag: xpath = f"//{parent_tag}/{tag_name}" conditions = [] if text_contains: conditions.append(f"contains(text(), '{text_contains}')") if position: conditions.append(f"position()={position}") if conditions: xpath += f"[{''.join(conditions)}]" return xpath # 示例调用 print(generate_xpath("title", text_contains="Python")) # 输出://title[contains(text(), 'Python')] print(generate_xpath("book", position=2)) # 输出://book[position()=2]
内容的提问来源于stack exchange,提问作者Kakashi
相关产品推荐
相关产品推荐

