Python处理XML文件:子节点访问与focus、type属性提取失败问题求助
Let's break down what's going wrong here and fix your XML data extraction code step by step. I can see you're trying to pull specific attributes and text from your XML structure, but the focus and type values from the ANNOTATION node aren't making it into your lists—let's dig into why.
First, Spot the Critical Issues in Your Current Code
Looking at your nested loops, there are a few key mistakes that are blocking the ANNOTATION data from being captured:
Incorrect Loop Nesting for
SUB-QUESTIONS
You're trying to accessSUB-QUESTIONSas a child ofMESSAGE, which almost certainly doesn't match your XML structure (based on your target field description).SUB-QUESTIONSis likely a direct child ofNLM-QUESTION, not nested underMESSAGE. Your current code checks ifsubchild.tag(which isMESSAGE) equalsSUB-QUESTIONS—this condition will never be true, so the code responsible for capturingfocusandtypenever runs properly.Variable Name Typo
You havefor subhild1 in subchild:—that's a typo (subhild1should besubchild1). This could cause runtime errors or make the loop target the wrong nodes entirely.Overly Deep & Unnecessary Nested Loops
You're looping throughsubchild3thensubchild4to findANNOTATION, but ifANNOTATIONis a direct child ofSUB-QUESTION, this extra nesting is redundant and risks missing the node.
Refactored Code to Fix These Issues
Here's a cleaned-up version that correctly targets the nodes you need, with comments explaining key changes:
for child in root: if child.tag == "NLM-QUESTION": # Extract questionid from NLM-QUESTION attributes for name, value in child.attrib.items(): if name == "questionid": quesid.append(value) # Loop through direct children of NLM-QUESTION to find MESSAGE and SUB-QUESTIONS for subchild in child: if subchild.tag == "MESSAGE": message.append(subchild.text) # Handle SUB-QUESTIONS directly under NLM-QUESTION elif subchild.tag == "SUB-QUESTIONS": # Loop through each SUB-QUESTION inside SUB-QUESTIONS for sub_question in subchild: if sub_question.tag == "SUB-QUESTION": # Extract subqid (subquestionid) for name, value in sub_question.attrib.items(): if name == "subqid": subquestionId.append(value) # Find ANNOTATION directly under SUB-QUESTION (adjust if nested deeper) for annotation_node in sub_question: if annotation_node.tag == "ANNOTATION": # Extract focus and type attributes (cleaner with get()) focus_val = annotation_node.get("focus") type_val = annotation_node.get("type") if focus_val: focus.append(focus_val) if type_val: type.append(type_val)
Additional Troubleshooting Tips
If the refactored code still doesn't capture focus and type, try these checks:
- Print Node Structures: Add
print(sub_question.tag)orprint(list(sub_question))inside theSUB-QUESTIONloop to confirmANNOTATIONis actually a child of that node. Watch out for XML namespaces (like{http://example.com}ANNOTATION) that might make the tag name look different than expected. - Verify Attribute Existence: Double-check that your
ANNOTATIONnodes actually havefocusandtypeattributes in the raw XML—missing attributes will result in no values being added. - Simplify Attribute Checks: Using
node.get("attribute_name")is more concise than looping through all attributes, and it returnsNoneif the attribute doesn't exist (which we handle with theif focus_valchecks).
内容的提问来源于stack exchange,提问作者Ayush Singh

