如何用Python XPath findall解析XML,获取指定节点的所有子元素?
解决XML中获取cpp:ifndef和cpp:define节点及其子元素的问题
问题描述
现有如下XML片段,包含
cpp:ifndef与cpp:define节点:... <cpp:ifndef pos:start="70:1" pos:end="70:22">#<cpp:directive pos:start="70:2" pos:end="70:7">ifndef</cpp:directive> <name pos:start="70:9" pos:end="70:22">INC_FREERTOS_H</name></cpp:ifndef> <cpp:define pos:start="71:1" pos:end="71:22">#<cpp:directive pos:start="71:2" pos:end="71:7">define</cpp:directive> <cpp:macro pos:start="71:9" pos:end="71:22"><name pos:start="71:9" pos:end="71:22">INC_FREERTOS_H</name></cpp:macro></cpp:define> ...需要获取
cpp:ifndef与cpp:define之间的所有元素,同时提取这两个节点内的cpp:directive、cpp:macro、name等子元素。但目前使用以下代码仅能获取节点的pos:start、pos:end属性,无法获取子元素:root.findall(".//cpp:ifndef", namespaces) root.findall(".//cpp:define", namespaces)
解决方案
你已经通过findall拿到了目标节点,接下来只需遍历这些节点,用find()或findall()提取子元素,同时可处理节点内的文本内容(比如开头的#)。
示例代码
# 假设已定义命名空间映射,比如namespaces={'cpp': 'http://example.com/cpp', 'pos': 'http://example.com/pos'} # 处理cpp:ifndef节点 ifndef_nodes = root.findall(".//cpp:ifndef", namespaces) for node in ifndef_nodes: # 获取pos属性(需使用带命名空间的完整键名) start_pos = node.get('{http://example.com/pos}start') end_pos = node.get('{http://example.com/pos}end') # 获取cpp:directive子元素及属性 directive = node.find('./cpp:directive', namespaces) directive_content = directive.text.strip() if directive else '' directive_start = directive.get('{http://example.com/pos}start') if directive else '' # 获取name子元素及属性 name = node.find('./name', namespaces) name_content = name.text.strip() if name else '' name_start = name.get('{http://example.com/pos}start') if name else '' # 获取节点内的前置文本(比如#) prefix_text = node.text.strip() if node.text else '' # 输出结果 print(f"ifndef节点: 起始位置{start_pos}, 结束位置{end_pos}") print(f" 前置文本: {prefix_text}") print(f" 指令: {directive_content} (位置{directive_start})") print(f" 宏名: {name_content} (位置{name_start})") # 处理cpp:define节点 define_nodes = root.findall(".//cpp:define", namespaces) for node in define_nodes: start_pos = node.get('{http://example.com/pos}start') end_pos = node.get('{http://example.com/pos}end') directive = node.find('./cpp:directive', namespaces) directive_content = directive.text.strip() if directive else '' # 嵌套获取cpp:macro下的name元素 macro = node.find('./cpp:macro', namespaces) macro_name_content = '' if macro: macro_name = macro.find('./name', namespaces) macro_name_content = macro_name.text.strip() if macro_name else '' prefix_text = node.text.strip() if node.text else '' print(f"\ndefine节点: 起始位置{start_pos}, 结束位置{end_pos}") print(f" 前置文本: {prefix_text}") print(f" 指令: {directive_content}") print(f" 宏名: {macro_name_content}")
获取两个节点之间的元素
如果需要收集cpp:ifndef和cpp:define之间的所有元素,可以通过遍历节点的兄弟节点实现:
ifndef_node = root.find(".//cpp:ifndef", namespaces) if ifndef_node: current_node = ifndef_node.next_sibling between_elements = [] while current_node is not None: # 遇到cpp:define节点时停止遍历 if current_node.tag == f"{{{namespaces['cpp']}}}define": break # 过滤空白文本节点,收集有效元素 if current_node.nodeType != current_node.TEXT_NODE or current_node.text.strip(): between_elements.append(current_node) current_node = current_node.next_sibling # 输出中间元素 for elem in between_elements: elem_tag = elem.tag.split('}')[-1] # 提取不带命名空间的标签名 print(f"中间元素: {elem_tag}, 内容: {elem.text.strip()}")
内容的提问来源于stack exchange,提问作者lebhero
相关产品推荐
相关产品推荐

