如何在XML中依据ImageType条件筛选对应的BinaryImage元素
问题描述
需要从以下XML文档中提取ImageType为fullImage的Image节点下的BinaryImage元素:
<?xml version="1.0" encoding="UTF-8"?> <soapenv:Envelope xmlns:soapenv="http://schemas.xmlsoap.org/soap/envelope/"> <soapenv:Header /> <soapenv:Body> <Instation xmlns="http://ffsf.us.com/schema_1-2" SchemaVersion="1.2"> <ImageArray> <Image> <InstanceID>5216</InstanceID> <TimeStamp>2022-12-01T10:34:24.499Z</TimeStamp> <LaneID>0</LaneID> <ImageType>fullImage</ImageType> <ImageFormat>jpeg</ImageFormat> <BinaryImage>abcd</BinaryImage> </Image> <Image> <InstanceID>5216</InstanceID> <TimeStamp>2022-12-01T10:34:24.499Z</TimeStamp> <LaneID>0</LaneID> <ImageType>Patch</ImageType> <ImageFormat>jpeg</ImageFormat> <BinaryImage>abcd</BinaryImage> </Image> </ImageArray> </Instation> </soapenv:Body> </soapenv:Envelope>
尝试了以下代码:
root.findall(".//{http://ffsf.us.com/schema_1-2}Image[contains(@ImageType,'fullImage')]") root.xpath(".//{http://ffsf.us.com/schema_1-2}Image[contains(@ImageType,'fullImage')]") root.xpath(".//{http://ffsf.us.com/schema_1-2}BinaryImage[contains(@ImageType,'fullImage')]") root.xpath(".//{http://ffsf.us.com/schema_1-2}BinaryImage[@ImageType='fullImage']") root.xpath(".//{http://ffsf.us.com/schema_1-2}Image[@ImageType='fullImage']")
但均出现错误:
lxml.etree.XPathEvalError: Invalid expression
SyntaxError: invalid predicate
错误原因
核心错误是把ImageType当成了Image节点的属性,但实际上ImageType是Image的子元素,不是属性。XPath中@符号专门用来选取节点属性,你需要的是匹配子元素的文本值,因此不能用@ImageType的写法。
解决方案
方法1:使用标准库xml.etree.ElementTree的findall
import xml.etree.ElementTree as ET tree = ET.parse("your_xml_file.xml") root = tree.getroot() # 定义命名空间映射 ns = {"ns": "http://ffsf.us.com/schema_1-2"} # 先筛选出ImageType为fullImage的Image节点,再提取BinaryImage full_image_nodes = root.findall(".//ns:Image[ns:ImageType='fullImage']", ns) for img_node in full_image_nodes: binary_image = img_node.find("ns:BinaryImage", ns) print(binary_image.text)
方法2:使用lxml的xpath(更灵活)
可以先注册命名空间,让XPath表达式更易读:
from lxml import etree tree = etree.parse("your_xml_file.xml") root = tree.getroot() # 注册命名空间 ns = {"ns": "http://ffsf.us.com/schema_1-2"} # 直接提取目标BinaryImage的文本内容 binary_images = root.xpath(".//ns:Image[ns:ImageType='fullImage']/ns:BinaryImage/text()", namespaces=ns) for bi in binary_images: print(bi)
如果不想注册命名空间,也可以直接使用完整的命名空间URI:
binary_images = root.xpath(".//{http://ffsf.us.com/schema_1-2}Image[{http://ffsf.us.com/schema_1-2}ImageType='fullImage']/{http://ffsf.us.com/schema_1-2}BinaryImage/text()")
补充说明
如果需要模糊匹配(比如ImageType文本包含fullImage),可以改用contains函数:
# lxml xpath写法 binary_images = root.xpath(".//ns:Image[contains(ns:ImageType/text(), 'fullImage')]/ns:BinaryImage/text()", namespaces=ns)
内容的提问来源于stack exchange,提问作者Vishal Singh
相关产品推荐
相关产品推荐

