You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中将SOAP接口返回的XML文本转换为Dataframe?

如何将转义后的SOAP响应文本转换为DataFrame?

问题场景

通过Python的requests库调用SOAP API后,返回的response.text中包含双重转义的XML数据:外层是正常的SOAP信封结构,但<return>标签内的内容被多次转义(比如<变成了&amp;lt;),需要先处理转义,再解析XML,最终转换为Pandas DataFrame。

解决步骤

1. 还原转义的XML内容

首先需要将<return>标签内的转义文本还原为正常XML。可以使用Python内置的html模块的unescape()方法,对内容进行两次解码(因为内容被双重转义)。

2. 解析XML提取数据

使用xml.etree.ElementTree解析还原后的XML,定位到ItemLocalidad节点,提取每个节点下的字段值。

3. 转换为DataFrame

将提取到的字段数据整理成列表字典的形式,再通过Pandas的DataFrame()方法转换为表格。

完整代码示例

import requests
import html
import xml.etree.ElementTree as ET
import pandas as pd

# 原SOAP请求代码
url = "https://url.com"
payload = """\n<soapenv:Envelope xmlns:soapenv="http://schemas.xmlsoap.org/soap/envelope/" xmlns:util="http://url.com/">
 <soapenv:Header/>
  <soapenv:Body>
   <util:localidadesListar>
    <usuario>user</usuario>
    <clave>password</clave>
   </util:localidadesListar>
  </soapenv:Body>
</soapenv:Envelope> 
"""
headers = {'Content-Type': 'text/xml; charset=utf-8'}
response = requests.request("POST", url, data=payload, headers=headers, verify=False)

# 处理响应数据
# 1. 解析外层SOAP响应,提取<return>标签内的内容
root = ET.fromstring(response.text)
# 注意:命名空间需与API返回的实际命名空间一致
return_content = root.find('.//{http://url.com/}return').text.strip()

# 2. 双重解码转义内容,还原为正常XML
decoded_xml = html.unescape(html.unescape(return_content))

# 3. 解析还原后的XML,提取ItemLocalidad节点数据
resp_root = ET.fromstring(decoded_xml)
items = resp_root.findall('.//ItemLocalidad')

# 4. 整理数据并转换为DataFrame
data_list = []
for item in items:
    item_dict = {
        'CodigoLocalidad': item.find('CodigoLocalidad').text or '',
        'DescripcionLocalidad': item.find('DescripcionLocalidad').text or '',
        'AbreviaturaLocalidad': item.find('AbreviaturaLocalidad').text or '',
        'DepartamentoLocalidad': item.find('DepartamentoLocalidad').text or ''
    }
    data_list.append(item_dict)

df = pd.DataFrame(data_list)
print(df)

关键注意事项

  • 命名空间匹配:外层SOAP响应的命名空间(示例中{http://url.com/})必须和API返回的实际命名空间完全一致,否则无法正确定位<return>标签。
  • 空值处理:代码中用or ''处理空标签场景,避免因节点无文本导致报错。
  • 依赖安装:若未安装Pandas,需先执行pip install pandas完成依赖配置。

内容的提问来源于stack exchange,提问作者Edjotace

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 11:52:23