如何使用XPath访问XML中深层嵌套的标签及其属性?
解决深层嵌套XML中Movie标签的XPath访问问题
首先得点明你代码里的核心问题:etree.tostring()返回的是XML的字节串表示,它不是可操作的Element对象,根本没有xpath()方法——你得到空列表的本质是用错了对象来调用XPath方法。下面一步步给你正确的解决方案:
1. 先正确解析XML,拿到可操作的Element对象
你需要先把XML内容解析成Element(或ElementTree)对象,而不是转成字符串。示例代码如下:
import xml.etree.ElementTree as etree # 你的XML内容 xml_content = '''<?xml version='1.0' encoding='utf8'?> <collection> <genre category="Action"> <decade years="1980s"> <movie favorite="True" title="Indiana Jones: The raiders of the lost Ark"> <format multiple="No">DVD</format> <year>1981</year> <rating>PG</rating> <description> 'Archaeologist and adventurer Indiana Jones is hired by the U.S. government to find the Ark of the Covenant before the Nazis.' </description> </movie> </decade> </genre> </collection>''' # 解析XML,得到根节点Element对象 root = etree.fromstring(xml_content)
2. 用正确的XPath定位Movie标签
因为Movie标签嵌套在collection > genre > decade层级下,有两种灵活的定位方式:
- 绝对路径定位:从根节点开始写完整路径,适合结构固定的XML:
# 绝对路径定位movie标签 movies = root.xpath('/collection/genre/decade/movie') - 全局匹配定位:用
//匹配任意层级的movie标签,适合XML结构有变动或者要找所有movie的场景:# 匹配所有层级的movie标签 movies = root.xpath('//movie')
这时候movies就会返回包含目标Movie元素的列表,不会再是空的了。
3. 访问Movie标签的属性
拿到Movie元素后,有两种方式获取它的属性:
- 通过Element的attrib字典:
if movies: target_movie = movies[0] # 获取favorite属性 print("是否为收藏:", target_movie.attrib.get('favorite')) # 获取title属性 print("电影标题:", target_movie.attrib.get('title')) - 直接用XPath提取属性:不用先获取元素,直接通过XPath表达式拿到属性值:
# 直接获取所有movie的title属性 movie_titles = root.xpath('//movie/@title') print("电影标题:", movie_titles[0]) # 筛选出favorite属性为True的movie标签 favorite_movies = root.xpath('//movie[@favorite="True"]')
内容的提问来源于stack exchange,提问作者Chang Zhao
相关产品推荐
相关产品推荐

