You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何传递XPath表达式及通过标签名生成XPath?

在Python中用XPath解析本地XML并生成XPath表达式

一、用XPath提取本地XML数据

Python里处理XPath最常用的是lxml库(对XPath语法支持更完整,比标准库的xml.etree.ElementTree功能更强),操作步骤如下:

  1. 安装lxml依赖
pip install lxml
  1. 加载本地XML并执行XPath查询
    假设你的XML文件为data.xml,内容示例:
<library>
    <book category="tech">
        <title>Python核心编程</title>
        <author>王五</author>
    </book>
    <book category="literature">
        <title>百年孤独</title>
        <author>加西亚·马尔克斯</author>
    </book>
</library>

代码示例:

from lxml import etree

# 加载本地XML文件
tree = etree.parse("data.xml")

# 提取所有book节点下的title文本
titles = tree.xpath("//book/title/text()")
for title in titles:
    print(title)

# 提取category为tech的book的author
tech_authors = tree.xpath("//book[@category='tech']/author/text()")
print(tech_authors)

如果想用标准库xml.etree.ElementTree(仅支持XPath 1.0子集),代码如下:

import xml.etree.ElementTree as ET

tree = ET.parse("data.xml")
root = tree.getroot()

# 查找所有title元素
titles = root.findall(".//title")
for title in titles:
    print(title.text)

二、通过标签名生成XPath表达式

可以写一个轻量化函数,根据传入的标签名和可选参数生成不同场景的XPath:

基础版:支持父节点、属性匹配

def generate_xpath(tag_name, parent_tag=None, attr_match=None):
    # 默认匹配所有该标签
    xpath = f"//{tag_name}"
    # 限定父节点路径
    if parent_tag:
        xpath = f"//{parent_tag}/{tag_name}"
    # 添加属性匹配条件
    if attr_match:
        # attr_match为字典,比如{"category": "tech"}
        attr_str = " and ".join([f"@{k}='{v}'" for k, v in attr_match.items()])
        xpath = f"{xpath}[{attr_str}]"
    return xpath

# 示例调用
print(generate_xpath("title"))  # 输出://title
print(generate_xpath("title", parent_tag="book"))  # 输出://book/title
print(generate_xpath("book", attr_match={"category": "literature"}))  # 输出://book[@category='literature']

进阶版:支持文本匹配、位置筛选

def generate_xpath(tag_name, parent_tag=None, text_contains=None, position=None):
    xpath = f"//{tag_name}"
    if parent_tag:
        xpath = f"//{parent_tag}/{tag_name}"
    
    conditions = []
    if text_contains:
        conditions.append(f"contains(text(), '{text_contains}')")
    if position:
        conditions.append(f"position()={position}")
    
    if conditions:
        xpath += f"[{''.join(conditions)}]"
    return xpath

# 示例调用
print(generate_xpath("title", text_contains="Python"))  # 输出://title[contains(text(), 'Python')]
print(generate_xpath("book", position=2))  # 输出://book[position()=2]

内容的提问来源于stack exchange,提问作者Kakashi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 23:10:59