You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在XML中依据ImageType条件筛选对应的BinaryImage元素

问题描述

需要从以下XML文档中提取ImageType为fullImage的Image节点下的BinaryImage元素:

<?xml version="1.0" encoding="UTF-8"?>
<soapenv:Envelope xmlns:soapenv="http://schemas.xmlsoap.org/soap/envelope/">
   <soapenv:Header />
   <soapenv:Body>
      <Instation xmlns="http://ffsf.us.com/schema_1-2" SchemaVersion="1.2">
         <ImageArray>
            <Image>
               <InstanceID>5216</InstanceID>
               <TimeStamp>2022-12-01T10:34:24.499Z</TimeStamp>
               <LaneID>0</LaneID>
               <ImageType>fullImage</ImageType>
               <ImageFormat>jpeg</ImageFormat>
               <BinaryImage>abcd</BinaryImage>
            </Image>
            <Image>
               <InstanceID>5216</InstanceID>
               <TimeStamp>2022-12-01T10:34:24.499Z</TimeStamp>
               <LaneID>0</LaneID>
               <ImageType>Patch</ImageType>
               <ImageFormat>jpeg</ImageFormat>
               <BinaryImage>abcd</BinaryImage>
            </Image>
         </ImageArray>
      </Instation>
   </soapenv:Body>
</soapenv:Envelope>

尝试了以下代码:

root.findall(".//{http://ffsf.us.com/schema_1-2}Image[contains(@ImageType,'fullImage')]")

root.xpath(".//{http://ffsf.us.com/schema_1-2}Image[contains(@ImageType,'fullImage')]")

root.xpath(".//{http://ffsf.us.com/schema_1-2}BinaryImage[contains(@ImageType,'fullImage')]")

root.xpath(".//{http://ffsf.us.com/schema_1-2}BinaryImage[@ImageType='fullImage']")

root.xpath(".//{http://ffsf.us.com/schema_1-2}Image[@ImageType='fullImage']")

但均出现错误:

lxml.etree.XPathEvalError: Invalid expression
SyntaxError: invalid predicate

错误原因

核心错误是把ImageType当成了Image节点的属性,但实际上ImageType是Image的子元素,不是属性。XPath中@符号专门用来选取节点属性,你需要的是匹配子元素的文本值,因此不能用@ImageType的写法。

解决方案

方法1:使用标准库xml.etree.ElementTree的findall

import xml.etree.ElementTree as ET

tree = ET.parse("your_xml_file.xml")
root = tree.getroot()
# 定义命名空间映射
ns = {"ns": "http://ffsf.us.com/schema_1-2"}

# 先筛选出ImageType为fullImage的Image节点,再提取BinaryImage
full_image_nodes = root.findall(".//ns:Image[ns:ImageType='fullImage']", ns)
for img_node in full_image_nodes:
    binary_image = img_node.find("ns:BinaryImage", ns)
    print(binary_image.text)

方法2:使用lxml的xpath(更灵活)

可以先注册命名空间,让XPath表达式更易读:

from lxml import etree

tree = etree.parse("your_xml_file.xml")
root = tree.getroot()

# 注册命名空间
ns = {"ns": "http://ffsf.us.com/schema_1-2"}

# 直接提取目标BinaryImage的文本内容
binary_images = root.xpath(".//ns:Image[ns:ImageType='fullImage']/ns:BinaryImage/text()", namespaces=ns)
for bi in binary_images:
    print(bi)

如果不想注册命名空间,也可以直接使用完整的命名空间URI:

binary_images = root.xpath(".//{http://ffsf.us.com/schema_1-2}Image[{http://ffsf.us.com/schema_1-2}ImageType='fullImage']/{http://ffsf.us.com/schema_1-2}BinaryImage/text()")

补充说明

如果需要模糊匹配(比如ImageType文本包含fullImage),可以改用contains函数:

# lxml xpath写法
binary_images = root.xpath(".//ns:Image[contains(ns:ImageType/text(), 'fullImage')]/ns:BinaryImage/text()", namespaces=ns)

内容的提问来源于stack exchange,提问作者Vishal Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:40:17