如何从指定SOAP格式XML文档中提取数据?
Hey there! Let's walk through how to extract values from your SOAP XML document—since it uses namespaces, that's the most important detail to get right. First, let's start with a cleaned-up version of your XML (I filled in the missing Security namespace to make examples functional):
<s:Envelope xmlns:s="http://www.w3.org/2003/05/soap-envelope" xmlns:u="http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-utility-1.0.xsd">
<s:Header>return (<VsDebuggerCausalityData xmlns="http://schemas.microsoft.com/vstudio/diagnostics/servicemodelsink">uIDPo4tYpt6X40FEk+VSAe5mc8MAAAAAP497cBuXfk+uFIOY80O0iuLtIW56q7hLktgVYPhbnHMACQAA</VsDebuggerCausalityData>)
<o:Security s:mustUnderstand="1" xmlns:o="http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-secext-1.0.xsd">
</o:Security>
</s:Header>
</s:Envelope>
Method 1: Use Python's built-in xml.etree.ElementTree
This works great if you don't want to install extra libraries. The key is defining a namespace map to match the prefixes in your XML.
import xml.etree.ElementTree as ET # Your full XML content (adjust as needed) xml_content = '''<s:Envelope xmlns:s="http://www.w3.org/2003/05/soap-envelope" xmlns:u="http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-utility-1.0.xsd"> <s:Header> <VsDebuggerCausalityData xmlns="http://schemas.microsoft.com/vstudio/diagnostics/servicemodelsink">uIDPo4tYpt6X40FEk+VSAe5mc8MAAAAAP497cBuXfk+uFIOY80O0iuLtIW56q7hLktgVYPhbnHMACQAA</VsDebuggerCausalityData> <o:Security s:mustUnderstand="1" xmlns:o="http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-secext-1.0.xsd"> <o:UsernameToken> <o:Username>test_user</o:Username> </o:UsernameToken> </o:Security> </s:Header> </s:Envelope>''' # Parse the XML root = ET.fromstring(xml_content) # Define namespace mappings (match prefixes to their URIs) namespaces = { 's': 'http://www.w3.org/2003/05/soap-envelope', 'debug': 'http://schemas.microsoft.com/vstudio/diagnostics/servicemodelsink', 'o': 'http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-secext-1.0.xsd' } # Extract the VsDebuggerCausalityData value debug_data_element = root.find('.//debug:VsDebuggerCausalityData', namespaces) if debug_data_element: print(f"Debug Causality Data: {debug_data_element.text}") # Extract a value inside the Security element (e.g., Username) username_element = root.find('.//o:Username', namespaces) if username_element: print(f"Username: {username_element.text}") # Extract the s:mustUnderstand attribute from Security security_element = root.find('.//o:Security', namespaces) if security_element: # Attributes with namespaces need the full URI in curly braces must_understand = security_element.get('{http://www.w3.org/2003/05/soap-envelope}mustUnderstand') print(f"Security mustUnderstand: {must_understand}")
Method 2: Use lxml (more powerful XML processing)
If you need advanced XPath support or better performance, lxml is a great choice. First install it with pip install lxml.
from lxml import etree xml_content = '''<s:Envelope xmlns:s="http://www.w3.org/2003/05/soap-envelope" xmlns:u="http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-utility-1.0.xsd"> <s:Header> <VsDebuggerCausalityData xmlns="http://schemas.microsoft.com/vstudio/diagnostics/servicemodelsink">uIDPo4tYpt6X40FEk+VSAe5mc8MAAAAAP497cBuXfk+uFIOY80O0iuLtIW56q7hLktgVYPhbnHMACQAA</VsDebuggerCausalityData> <o:Security s:mustUnderstand="1" xmlns:o="http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-secext-1.0.xsd"> <o:UsernameToken> <o:Username>test_user</o:Username> </o:UsernameToken> </o:Security> </s:Header> </s:Envelope>''' # Parse the XML root = etree.fromstring(xml_content) # Extract debug data with XPath (specify namespace directly in the query) debug_data = root.xpath('//debug:VsDebuggerCausalityData/text()', namespaces={'debug': 'http://schemas.microsoft.com/vstudio/diagnostics/servicemodelsink'}) if debug_data: print(f"Debug Causality Data: {debug_data[0]}") # Extract Username username = root.xpath('//o:Username/text()', namespaces={'o': 'http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-secext-1.0.xsd'}) if username: print(f"Username: {username[0]}") # Extract the mustUnderstand attribute must_understand = root.xpath('//o:Security/@s:mustUnderstand', namespaces={'o': 'http://docs.oasis-open.org/wss/2004/01/oasis-200401-wss-wssecurity-secext-1.0.xsd', 's': 'http://www.w3.org/2003/05/soap-envelope'}) if must_understand: print(f"Security mustUnderstand: {must_understand[0]}")
Key Things to Remember
- Namespaces are non-negotiable: Every prefixed element/attribute (like
s:oro:) and elements with a default namespace (likeVsDebuggerCausalityData) requires you to map its URI to a prefix in your queries. Without this, your parser won't find the elements. - Attribute namespaces: Attributes with a prefix (like
s:mustUnderstand) belong to the prefix's namespace, so you need to include that in your query. - Valid XML first: Make sure your full XML is well-formed (no missing tags or truncated namespace URIs) before trying to parse it—broken XML will throw errors.
内容的提问来源于stack exchange,提问作者MindGame

