使用lxml解析带SBE命名空间的XML,获取sbe:messageSchema及字段
Got it, let's break down what's going on here and fix your lxml parsing issue with the SBE XML schema.
First, a quick clarification: The root node of your XML is the <sbe:messageSchema> element itself. When you were traversing the root's child nodes, the description and byteOrder you saw are attributes of this root node—not child nodes. The actual child nodes you're looking for are the <sbe:message> elements nested inside it.
The core problem here is XML namespace handling. lxml doesn't automatically recognize prefixes like sbe: from your XML; you need to explicitly map the prefix to its underlying namespace URI so the parser knows how to match elements.
Here's a step-by-step solution with code:
1. Define the Namespace Mapping
First, create a dictionary that maps the sbe prefix to its official URI from your XML: http://www.fixprotocol.org/ns/simple/1.0. This tells lxml what the sbe: prefix actually refers to.
2. Parse the XML and Access Elements
Use lxml.etree to parse your XML, then use iterfind (or xpath) with the namespace mapping to locate the elements you need.
from lxml import etree # Replace this with your actual XML content or file path xml_content = """<sbe:messageSchema xmlns:sbe="http://www.fixprotocol.org/ns/simple/1.0" description="something" byteOrder="littleEndian"> <sbe:message name="DummyMsg" id="99" description="Placeholder message. Uses otherwise unused enums and composites so sbe compiles them."> <field name="msgType" id="1" type="MsgType" /> <field name="minimumSbeSchemaVersion" id="2" type="MinimumSbeSchemaVersion"/> </sbe:message> </sbe:messageSchema>""" # Parse the XML string (use etree.parse() for files) root = etree.fromstring(xml_content) # Define the namespace mapping ns_map = {"sbe": "http://www.fixprotocol.org/ns/simple/1.0"} # Access the <sbe:messageSchema> attributes (since root is this element) print("=== messageSchema Details ===") print(f"Description: {root.get('description')}") print(f"Byte Order: {root.get('byteOrder')}") # Find all <sbe:message> elements under the root messages = root.iterfind("./sbe:message", namespaces=ns_map) for msg in messages: print("\n=== Message Details ===") print(f"Name: {msg.get('name')}") print(f"ID: {msg.get('id')}") print(f"Description: {msg.get('description')}") # Find all <field> elements inside the message (no prefix needed here) fields = msg.iterfind("./field") for field in fields: print(f"\nField:") print(f" Name: {field.get('name')}") print(f" ID: {field.get('id')}") print(f" Type: {field.get('type')}")
Why This Works
- Namespace Mapping: The
ns_mapdictionary links thesbeprefix to its URI, which is the true identifier for elements in this namespace (prefixes are just shortcuts). - iterfind with Namespaces: By passing
namespaces=ns_maptoiterfind, lxml correctly interprets thesbe:prefix in your query. - Unprefixed Elements: The
<field>elements don't use thesbeprefix, so you can query them directly without any namespace reference.
If your XML has multiple <sbe:messageSchema> elements (e.g., wrapped in a parent root), you can adjust the query to find them like this:
# Find all messageSchema elements in the document all_message_schemas = root.iterfind("//sbe:messageSchema", namespaces=ns_map) for schema in all_message_schemas: # Process each schema here pass
内容的提问来源于stack exchange,提问作者shrishinde

