从API XML响应构建Pandas DataFrame:遍历Results节点问题
遍历XML的Results节点并转换为Pandas DataFrame
问题描述
调用API接口得到XML响应,需要遍历其中的Results节点,并将所有数据存储为pandas DataFrame。已通过链式调用.find()方法获取到单条数据,但不清楚如何循环遍历Body中的所有Results块。使用环境为Windows系统下Jupyter中的Python 3.7+。
已尝试代码
import pandas as pd from bs4 import BeautifulSoup import xml.etree.ElementTree as ET soup = BeautifulSoup(soap_response.text, "xml") # print(soup.prettify()) objectid_field = soup.find('Results').find('ObjectID').text customerkey_field = soup.find('Results').find('CustomerKey').text name_field = soup.find('Results').find('Name').text issendable_field = name_field = soup.find('Results').find('IsSendable').text sendablesubscribe_field = soup.find('Results').find('SendableSubscriberField').text # for de in soup: # de_name = soup.find('Results').find('Name').text # print(de_name) # test_df = pd.read_xml(soup, # xpath="//Results", # namespaces={""})
示例XML数据结构
<?xml version="1.0" encoding="utf-8"?> <soap:Envelope xmlns:soap="http://www.w3.org/2003/soap-envelope" xmlns:xsi="http://www.w3.org/2001/XMLSchema" xmlns:xsd="http://www.w3.org/2003/XMLSchema" xmlns:wsa="http://schemas.xmlsoap.org/ws/2004/08/addressing" xmlns:wsse="http://docs.oasis-open.org/wss/2004/01/oasis-201-wss-wssecurity-secext-1.0.xsd" xmlns:wsu="http://docs.oasis-open.org/wss/2004/01/oasis-201-wss-security-1.0.xsd"> <env:Header xmlns:env="http://www.w3.org/2003/05/soap-envelope"> <wsa:Action>RetrieveResponse</wsa:Action> <wsa:MessageID>urn:uuid:1234</wsa:MessageID> <wsa:RelatesTo>urn:uuid:1234</wsa:RelatesTo> <wsa:To>http://schemas.xmlsoap.org/ws/2004/08/dressing/role/anonymous</wsa:To> <wsse:Security> <wsu:Timestamp wsu:Id="Timestamp-1234"> <wsu:Created>2021-11-07T13:10:54Z</wsu:Created> <wsu:Expires>2021-11-07T13:15:54Z</wsu:Expires> </wsu:Timestamp> </wsse:Security> </env:Header> <soap:Body> <RetrieveResponseMsg xmlns="http://partnerAPI"> <OverallStatus>OK</OverallStatus> <RequestID>f9876</RequestID> <Results xsi:type="Data"> <PartnerKey xsi:nil="true" /> <ObjectID>Object1</ObjectID> <CustomerKey>Customer1</CustomerKey> <Name>Test1</Name> <IsSendable>true</IsSendable> <SendableSubscriberField> <Name>_Something1</Name> </SendableSubscriberField> </Results> <Results xsi:type="Data"> <PartnerKey xsi:nil="true" /> <ObjectID>Object2</ObjectID> <CustomerKey>Customer2</CustomerKey> <Name>Name2</Name> <IsSendable>true</IsSendable> <SendableSubscriberField> <Name>_Something2</Name> </SendableSubscriberField> </Results> <Results xsi:type="Data"> <PartnerKey xsi:nil="true" /> <ObjectID>Object3</ObjectID> <CustomerKey>AnotherKey</CustomerKey> <Name>Something3</Name> <IsSendable>false</IsSendable> </Results> </RetrieveResponseMsg> </soap:Body> </soap:Envelope>
解决方案
方法一:使用BeautifulSoup遍历所有Results节点
通过soup.find_all('Results')获取所有Results节点,循环遍历每个节点提取字段,同时处理部分节点可能缺失的字段(比如第三个Results没有SendableSubscriberField)。
import pandas as pd from bs4 import BeautifulSoup # 假设soap_response是你的API响应对象 soup = BeautifulSoup(soap_response.text, "xml") # 初始化存储数据的列表 data_list = [] # 获取所有Results节点 all_results = soup.find_all('Results') for result in all_results: # 提取字段,处理可能不存在的节点 record = { 'ObjectID': result.find('ObjectID').text if result.find('ObjectID') else None, 'CustomerKey': result.find('CustomerKey').text if result.find('CustomerKey') else None, 'Name': result.find('Name').text if result.find('Name') else None, 'IsSendable': result.find('IsSendable').text if result.find('IsSendable') else None, # 处理嵌套的SendableSubscriberField节点 'SendableSubscriberField': result.find('SendableSubscriberField').find('Name').text if result.find('SendableSubscriberField') else None } data_list.append(record) # 转换为DataFrame df = pd.DataFrame(data_list) print(df)
方法二:使用pandas.read_xml直接解析(更简洁)
利用pandas内置的read_xml方法,通过指定命名空间和XPath定位所有Results节点,再处理嵌套字段。
import pandas as pd # 直接从API响应文本解析XML df = pd.read_xml( soap_response.text, xpath="//ns:Results", namespaces={"ns": "http://partnerAPI"} ) # 提取嵌套的SendableSubscriberField中的Name值 df['SendableSubscriberField'] = df['SendableSubscriberField'].apply( lambda x: x['Name'] if isinstance(x, dict) else None ) print(df)
说明
- BeautifulSoup方法:灵活性更高,适合处理节点结构不一致的XML,能轻松判断并处理缺失字段。
- pandas.read_xml方法:代码更简洁,适合结构规整的XML,无需手动循环遍历,但需要正确处理命名空间和嵌套字段。
内容的提问来源于stack exchange,提问作者Carson Whitley
相关产品推荐
相关产品推荐

