You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python XPath findall解析XML,获取指定节点的所有子元素?

解决XML中获取cpp:ifndef和cpp:define节点及其子元素的问题

问题描述

现有如下XML片段,包含cpp:ifndef与cpp:define节点:

...
<cpp:ifndef pos:start="70:1" pos:end="70:22">#<cpp:directive pos:start="70:2" pos:end="70:7">ifndef</cpp:directive> <name pos:start="70:9" pos:end="70:22">INC_FREERTOS_H</name></cpp:ifndef>
<cpp:define pos:start="71:1" pos:end="71:22">#<cpp:directive pos:start="71:2" pos:end="71:7">define</cpp:directive> <cpp:macro pos:start="71:9" pos:end="71:22"><name pos:start="71:9" pos:end="71:22">INC_FREERTOS_H</name></cpp:macro></cpp:define>
...

需要获取cpp:ifndef与cpp:define之间的所有元素,同时提取这两个节点内的cpp:directive、cpp:macro、name等子元素。但目前使用以下代码仅能获取节点的pos:start、pos:end属性,无法获取子元素:

root.findall(".//cpp:ifndef", namespaces)
root.findall(".//cpp:define", namespaces)

解决方案

你已经通过findall拿到了目标节点,接下来只需遍历这些节点,用find()或findall()提取子元素,同时可处理节点内的文本内容(比如开头的#)。

示例代码

# 假设已定义命名空间映射,比如namespaces={'cpp': 'http://example.com/cpp', 'pos': 'http://example.com/pos'}

# 处理cpp:ifndef节点
ifndef_nodes = root.findall(".//cpp:ifndef", namespaces)
for node in ifndef_nodes:
    # 获取pos属性(需使用带命名空间的完整键名)
    start_pos = node.get('{http://example.com/pos}start')
    end_pos = node.get('{http://example.com/pos}end')
    
    # 获取cpp:directive子元素及属性
    directive = node.find('./cpp:directive', namespaces)
    directive_content = directive.text.strip() if directive else ''
    directive_start = directive.get('{http://example.com/pos}start') if directive else ''
    
    # 获取name子元素及属性
    name = node.find('./name', namespaces)
    name_content = name.text.strip() if name else ''
    name_start = name.get('{http://example.com/pos}start') if name else ''
    
    # 获取节点内的前置文本(比如#)
    prefix_text = node.text.strip() if node.text else ''
    
    # 输出结果
    print(f"ifndef节点: 起始位置{start_pos}, 结束位置{end_pos}")
    print(f"  前置文本: {prefix_text}")
    print(f"  指令: {directive_content} (位置{directive_start})")
    print(f"  宏名: {name_content} (位置{name_start})")

# 处理cpp:define节点
define_nodes = root.findall(".//cpp:define", namespaces)
for node in define_nodes:
    start_pos = node.get('{http://example.com/pos}start')
    end_pos = node.get('{http://example.com/pos}end')
    
    directive = node.find('./cpp:directive', namespaces)
    directive_content = directive.text.strip() if directive else ''
    
    # 嵌套获取cpp:macro下的name元素
    macro = node.find('./cpp:macro', namespaces)
    macro_name_content = ''
    if macro:
        macro_name = macro.find('./name', namespaces)
        macro_name_content = macro_name.text.strip() if macro_name else ''
    
    prefix_text = node.text.strip() if node.text else ''
    
    print(f"\ndefine节点: 起始位置{start_pos}, 结束位置{end_pos}")
    print(f"  前置文本: {prefix_text}")
    print(f"  指令: {directive_content}")
    print(f"  宏名: {macro_name_content}")

获取两个节点之间的元素

如果需要收集cpp:ifndef和cpp:define之间的所有元素,可以通过遍历节点的兄弟节点实现:

ifndef_node = root.find(".//cpp:ifndef", namespaces)
if ifndef_node:
    current_node = ifndef_node.next_sibling
    between_elements = []
    
    while current_node is not None:
        # 遇到cpp:define节点时停止遍历
        if current_node.tag == f"{{{namespaces['cpp']}}}define":
            break
        # 过滤空白文本节点,收集有效元素
        if current_node.nodeType != current_node.TEXT_NODE or current_node.text.strip():
            between_elements.append(current_node)
        current_node = current_node.next_sibling
    
    # 输出中间元素
    for elem in between_elements:
        elem_tag = elem.tag.split('}')[-1]  # 提取不带命名空间的标签名
        print(f"中间元素: {elem_tag}, 内容: {elem.text.strip()}")

内容的提问来源于stack exchange,提问作者lebhero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 08:02:03