You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理XML文件:子节点访问与focus、type属性提取失败问题求助

Troubleshooting XML Extraction: Focus & Type Attributes Not Being Captured

Let's break down what's going wrong here and fix your XML data extraction code step by step. I can see you're trying to pull specific attributes and text from your XML structure, but the focus and type values from the ANNOTATION node aren't making it into your lists—let's dig into why.

First, Spot the Critical Issues in Your Current Code

Looking at your nested loops, there are a few key mistakes that are blocking the ANNOTATION data from being captured:

  • Incorrect Loop Nesting for SUB-QUESTIONS
    You're trying to access SUB-QUESTIONS as a child of MESSAGE, which almost certainly doesn't match your XML structure (based on your target field description). SUB-QUESTIONS is likely a direct child of NLM-QUESTION, not nested under MESSAGE. Your current code checks if subchild.tag (which is MESSAGE) equals SUB-QUESTIONS—this condition will never be true, so the code responsible for capturing focus and type never runs properly.

  • Variable Name Typo
    You have for subhild1 in subchild:—that's a typo (subhild1 should be subchild1). This could cause runtime errors or make the loop target the wrong nodes entirely.

  • Overly Deep & Unnecessary Nested Loops
    You're looping through subchild3 then subchild4 to find ANNOTATION, but if ANNOTATION is a direct child of SUB-QUESTION, this extra nesting is redundant and risks missing the node.

Refactored Code to Fix These Issues

Here's a cleaned-up version that correctly targets the nodes you need, with comments explaining key changes:

for child in root:
    if child.tag == "NLM-QUESTION":
        # Extract questionid from NLM-QUESTION attributes
        for name, value in child.attrib.items():
            if name == "questionid":
                quesid.append(value)
        
        # Loop through direct children of NLM-QUESTION to find MESSAGE and SUB-QUESTIONS
        for subchild in child:
            if subchild.tag == "MESSAGE":
                message.append(subchild.text)
            
            # Handle SUB-QUESTIONS directly under NLM-QUESTION
            elif subchild.tag == "SUB-QUESTIONS":
                # Loop through each SUB-QUESTION inside SUB-QUESTIONS
                for sub_question in subchild:
                    if sub_question.tag == "SUB-QUESTION":
                        # Extract subqid (subquestionid)
                        for name, value in sub_question.attrib.items():
                            if name == "subqid":
                                subquestionId.append(value)
                        
                        # Find ANNOTATION directly under SUB-QUESTION (adjust if nested deeper)
                        for annotation_node in sub_question:
                            if annotation_node.tag == "ANNOTATION":
                                # Extract focus and type attributes (cleaner with get())
                                focus_val = annotation_node.get("focus")
                                type_val = annotation_node.get("type")
                                if focus_val:
                                    focus.append(focus_val)
                                if type_val:
                                    type.append(type_val)

Additional Troubleshooting Tips

If the refactored code still doesn't capture focus and type, try these checks:

  • Print Node Structures: Add print(sub_question.tag) or print(list(sub_question)) inside the SUB-QUESTION loop to confirm ANNOTATION is actually a child of that node. Watch out for XML namespaces (like {http://example.com}ANNOTATION) that might make the tag name look different than expected.
  • Verify Attribute Existence: Double-check that your ANNOTATION nodes actually have focus and type attributes in the raw XML—missing attributes will result in no values being added.
  • Simplify Attribute Checks: Using node.get("attribute_name") is more concise than looping through all attributes, and it returns None if the attribute doesn't exist (which we handle with the if focus_val checks).

内容的提问来源于stack exchange,提问作者Ayush Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 01:29:07