Python遍历XML标签所有子标签与文本(无需指定子标签名)及XML片段处理问询
Got it, let's walk through how to traverse all child tags and text content in your XML snippet without having to hardcode any child tag names. We'll use Python's built-in xml.etree.ElementTree module—no extra installs needed, which is perfect for quick, reliable XML processing.
First, let's recap your target XML snippet (formatted for readability):
<eligibility> <criteria> <textblock> Inclusion Criteria: - women undergoing cesarean section for any indication - literate in german language Exclusion Criteria: - history of keloids - previous transversal suprapubic scars - known patient hypersensitivity to any of the suture materials used in the protocol - a medical disorder that could affect wound healing (eg, diabetes mellitus, chronic corticosteroid use) </textblock> </criteria> <gender>Female</gender> <minimum_age>18 Years</minimum_age> <maximum_age>45 Years</maximum_age> <healthy_volunteers>No</healthy_volunteers> </eligibility>
Step 1: Parse the XML
First, we'll parse the XML content. You can read it from a file or pass it as a string directly. Here's how to handle a string:
import xml.etree.ElementTree as ET # Your XML content as a string xml_content = """ <eligibility> <criteria> <textblock> Inclusion Criteria: - women undergoing cesarean section for any indication - literate in german language Exclusion Criteria: - history of keloids - previous transversal suprapubic scars - known patient hypersensitivity to any of the suture materials used in the protocol - a medical disorder that could affect wound healing (eg, diabetes mellitus, chronic corticosteroid use) </textblock> </criteria> <gender>Female</gender> <minimum_age>18 Years</minimum_age> <maximum_age>45 Years</maximum_age> <healthy_volunteers>No</healthy_volunteers> </eligibility> """ # Parse the XML root = ET.fromstring(xml_content)
Step 2: Recursive Traversal Function
The key here is a recursive function that will dig into every node—whether it's a tag or text—without needing to know tag names ahead of time. We'll skip empty text nodes (like whitespace between tags) to keep things clean:
def traverse_xml(node, indent=0): # Print the current tag (if it's an element) if node.tag is not ET.Comment: print(f"{' '*indent}Tag: *{node.tag}*") # Process text content (skip empty whitespace) text = node.text.strip() if node.text else "" if text: print(f"{' '*indent}Text: {text}") # Recursively traverse all child nodes for child in node: traverse_xml(child, indent + 1)
Step 3: Run the Traversal
Call the function with the root element, and you'll get a clean breakdown of all tags and text:
traverse_xml(root)
Sample Output
Here's what you'll see when you run the code:
Tag: *eligibility* Tag: *criteria* Tag: *textblock* Text: Inclusion Criteria: - women undergoing cesarean section for any indication - literate in german language Exclusion Criteria: - history of keloids - previous transversal suprapubic scars - known patient hypersensitivity to any of the suture materials used in the protocol - a medical disorder that could affect wound healing (eg, diabetes mellitus, chronic corticosteroid use) Tag: *gender* Text: Female Tag: *minimum_age* Text: 18 Years Tag: *maximum_age* Text: 45 Years Tag: *healthy_volunteers* Text: No
Customization Tips
- If you want to collect the data instead of printing it, modify the function to append to a list or dictionary instead of using
print(). - For XML files instead of strings, use
tree = ET.parse("your_file.xml")thenroot = tree.getroot(). - If you need to handle namespaces, add
ET.register_namespace(prefix, uri)before parsing to avoid namespace prefixes being stripped.
内容的提问来源于stack exchange,提问作者Slowat_Kela

