You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的lxml解析带命名空间的在线XML,提取文本遇阻求助

解决lxml提取带命名空间XML表格文本的问题

问题描述

我尝试从在线XML文件的表格中提取特定文本(Prepare to write scripts、Write draft scripts、Produce final scripts),已经通过以下代码定位到表格元素,但无法进一步提取目标文本:

from lxml import etree
import requests

main_file = requests.get('https://training.gov.au/TrainingComponentFiles/CUA/CUAWRT601_R1.xml')
main_file.encoding = 'utf-8-sig'
root = etree.fromstring(main_file.content)
tables = root.xpath('//foo:table', namespaces={"foo": "http://www.authorit.com/xml/authorit"})

print(tables)

在XPATH工具中可以通过表达式//table[1]/tr/td[@width="2700"]/p[@id="4"][not(*)]/text()获取目标文本,但该表达式在本地代码中无效,需要可行的解决方法。

解决方案

核心原因:命名空间未全局匹配

本地代码已定义命名空间前缀foo,但后续XPath查询时未给所有元素添加该前缀——目标XML中所有元素(table、tr、td、p)都属于http://www.authorit.com/xml/authorit命名空间,必须显式指定前缀才能正确匹配。

修改后的可运行代码

直接在XPath中为所有元素添加foo:前缀,无需先单独获取tables列表:

from lxml import etree
import requests

main_file = requests.get('https://training.gov.au/TrainingComponentFiles/CUA/CUAWRT601_R1.xml')
main_file.encoding = 'utf-8-sig'
root = etree.fromstring(main_file.content)

# 为所有XML元素添加命名空间前缀,执行查询
target_texts = root.xpath(
    '//foo:table[1]/foo:tr/foo:td[@width="2700"]/foo:p[@id="4"][not(*)]/text()',
    namespaces={"foo": "http://www.authorit.com/xml/authorit"}
)

# 输出清理后的目标文本
for text in target_texts:
    print(text.strip())

代码说明

  1. 命名空间前缀:所有XML元素(table、tr、td、p)必须添加foo:前缀,与代码中定义的命名空间映射对应。
  2. 过滤逻辑:[not(*)]确保只提取无嵌套子元素的p标签文本,避免混入无关内容。
  3. 文本清理:用strip()去除文本前后空白字符,得到整洁的目标内容。

运行代码后将输出:

Prepare to write scripts
Write draft scripts
Produce final scripts

内容的提问来源于stack exchange,提问作者lucas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 00:05:28