如何用Python ElementTree解析带命名空间的XML并获取非自闭合元素值?
问题:解析带命名空间的XML,获取非自闭合
<itunes:subtitle>标签的值 我遇到的问题和类似的命名空间XML解析问题类似,但始终无法解决。我需要获取所有**非自闭合标签<itunes:subtitle>**的值,注意要排除自闭合的<itunes:subtitle/>标签。
示例XML(sample.xml)
<?xml version="1.0" encoding="UTF-8" standalone="yes"?> <rss version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:content="http://purl.org/rss/1.0/modules/content/"> <channel> <item> <title>A title</title> <itunes:subtitle/> <itunes:subtitle>A subtitle</itunes:subtitle> </item> <item> <title><![CDATA[Another title]]></title> <itunes:subtitle/> <itunes:subtitle>Yet another subtitlen</itunes:subtitle> </item> </channel> </rss>
我的Python代码
import xml.etree.ElementTree as ET with open('sample.xml', 'r', encoding='utf8') as f: tree = ET.parse(f) root = tree.getroot() for xml_item in root.iter('item'): namespaces = {'itunes': 'subtitle'} print(root.findall('itunes:subtitle', namespaces))
运行结果
[] []
问题分析与解决方案
- 命名空间映射错误:你定义的
namespaces字典把itunes映射成了subtitle,这完全错误。正确的映射应该是XML里xmlns:itunes对应的URL:http://www.itunes.com/dtds/podcast-1.0.dtd。 - 查找路径错误:遍历
item元素时,却用root.findall查找,而root是<rss>节点,直接找itunes:subtitle会找不到——这些标签在item内部,应该用当前的xml_item调用findall。 - 筛选非自闭合标签:自闭合标签的
text属性为None,而非自闭合标签有文本内容,可通过判断element.text是否非空来筛选。
修正后的代码
import xml.etree.ElementTree as ET with open('sample.xml', 'r', encoding='utf8') as f: tree = ET.parse(f) root = tree.getroot() # 正确的命名空间映射 namespaces = {'itunes': 'http://www.itunes.com/dtds/podcast-1.0.dtd'} for xml_item in root.iter('item'): # 在当前item下查找所有itunes:subtitle标签 subtitles = xml_item.findall('itunes:subtitle', namespaces) # 筛选出有有效文本的非自闭合标签 valid_subtitles = [sub.text.strip() for sub in subtitles if sub.text and sub.text.strip()] print(valid_subtitles)
修正后运行结果
['A subtitle'] ['Yet another subtitlen']
内容的提问来源于stack exchange,提问作者Arete
相关产品推荐
相关产品推荐

