Python SOAP请求无法从美海岸警卫队公开数据库获取数据
问题描述
尝试从美国海岸警卫队公开数据库获取船舶缺陷数据,该数据库提供SOAP访问方式及返回变量定义。参考官方示例修改SOAP代码后用Python实现,输入的Activity Number为有效编号,请求未报错但未返回预期有效数据,同时需要提取返回结果中的System、Description、FailureCauseLookupName、IsResolved等字段。
原代码如下:
import requests from xml.etree import ElementTree as ET import pandas as pd url=https://cgmix.uscg.mil/xml/PSIXData.asmx headers= {'content-type': 'application/soap+xml', 'SOAPAction': 'http://cgmix.uscg.mil/getVesselDeficiencies'} body= """<soap12:Envelope xmlns:xsi=http://www.w3.org/2001/XMLSchema-instance xmlns:xsd=http://www.w3.org/2001/XMLSchema xmlns:soap12=http://www.w3.org/2003/05/soap-envelope> <soap12:Body> <getVesselDeficienciesXMLString xmlns=http://cgmix.uscg.mil> <ActivityNumber>6692725</ActivityNumber> </getVesselDeficienciesXMLString> </soap12:Body> </soap12:Envelope>""" response = requests.post(url,data=body,headers=headers) print (response.content)
解决方案
1. 修复SOAP请求核心问题
原代码存在XML语法错误和SOAPAction不匹配的问题,导致服务器无法正确处理请求。以下是修正后的完整代码,包含数据提取逻辑:
import requests from xml.etree import ElementTree as ET import pandas as pd # 修正URL和请求头格式 url = "https://cgmix.uscg.mil/xml/PSIXData.asmx" headers = { 'Content-Type': 'application/soap+xml; charset=utf-8', 'SOAPAction': 'http://cgmix.uscg.mil/getVesselDeficienciesXMLString' } # 修复XML属性引号缺失问题 body = """<soap12:Envelope xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xsd="http://www.w3.org/2001/XMLSchema" xmlns:soap12="http://www.w3.org/2003/05/soap-envelope"> <soap12:Body> <getVesselDeficienciesXMLString xmlns="http://cgmix.uscg.mil"> <ActivityNumber>6692725</ActivityNumber> </getVesselDeficienciesXMLString> </soap12:Body> </soap12:Envelope>""" # 发送请求并处理响应 response = requests.post(url, data=body.encode('utf-8'), headers=headers) # 解析SOAP响应XML root = ET.fromstring(response.content) # 定义命名空间映射,适配响应结构 ns = { 'soap12': 'http://www.w3.org/2003/05/soap-envelope', 'cg': 'http://cgmix.uscg.mil' } # 提取返回的嵌套XML字符串(接口返回的缺陷数据是XML格式的字符串) result_xml_str = root.find('.//cg:getVesselDeficienciesXMLStringResult', ns).text if result_xml_str: # 解析嵌套的缺陷数据XML def_root = ET.fromstring(result_xml_str) # 提取目标字段 deficiency_list = [] # 遍历所有缺陷条目(节点名需匹配接口实际返回的结构,此处假设为Deficiency) for item in def_root.findall('.//Deficiency'): deficiency = { 'System': item.find('System').text if item.find('System') is not None else None, 'Description': item.find('Description').text if item.find('Description') is not None else None, 'FailureCauseLookupName': item.find('FailureCauseLookupName').text if item.find('FailureCauseLookupName') is not None else None, 'IsResolved': item.find('IsResolved').text if item.find('IsResolved') is not None else None } deficiency_list.append(deficiency) # 转为DataFrame方便查看和处理 df = pd.DataFrame(deficiency_list) print(df) else: print("未查询到对应Activity Number的缺陷数据")
2. 关键修正说明
- XML语法修复:原请求体中
xmlns:xsi=http://...这类写法缺少双引号,XML规范要求属性值必须用引号包裹,服务器无法解析无效XML。 - SOAPAction匹配:原
SOAPAction值为getVesselDeficiencies,但接口实际操作名为getVesselDeficienciesXMLString,必须完全匹配才能触发正确的接口方法。 - 嵌套XML解析:接口返回的结果是一个包含XML字符串的节点,需要先提取该字符串再二次解析,才能获取具体的缺陷数据。
- 空值处理:添加了字段空值判断,避免因节点缺失导致的报错。
3. 调试建议
- 先打印
response.text查看原始响应内容,确认是否返回错误提示或空结果,区分是请求问题还是数据本身不存在。 - 对比官方提供的SOAP请求示例,检查请求体结构、命名空间、参数名是否完全一致。
- 验证
ActivityNumber是否确实对应有缺陷记录,可通过官方网页端查询确认。
内容的提问来源于stack exchange,提问作者RachZ
相关产品推荐
相关产品推荐

