如何在Python的lxml中编写XPath检查兄弟节点关系并获取属性
解决XPath语法错误及获取目标属性值的方案
修正XPath语法
你之前的XPath报错是因为在谓词中连续嵌套属性判断的写法不符合XPath 1.0规范(lxml默认采用XPath 1.0),可以通过合并属性条件来修复:
方案1:合并code节点的属性判断
将code节点的两个属性条件放在同一个谓词内,用and连接:
.//outboundRelationship[@typeCode='SPRT'][relatedInvestigation/code[@code='2' and @codeSystem='2.16.840.1.113883.3.989.2.1.1.22']]/priorityNumber
方案2:通过兄弟节点关系精准定位
如果需要明确指定priorityNumber是符合条件的relatedInvestigation的兄弟节点,可使用兄弟轴定位:
# 从priorityNumber出发,检查后续兄弟节点是否符合条件 .//priorityNumber[../@typeCode='SPRT' and following-sibling::relatedInvestigation/code[@code='2' and @codeSystem='2.16.840.1.113883.3.989.2.1.1.22']] # 或者从符合条件的code节点反向遍历到priorityNumber .//relatedInvestigation/code[@code='2' and @codeSystem='2.16.840.1.113883.3.989.2.1.1.22']/../../priorityNumber
替代实现方法
方法1:分步查找(可读性更强)
避免复杂XPath,在Python代码中分步骤定位元素,便于调试和维护:
import lxml.etree as ET def xml_get_attrib_value(filepath, attribute): it = ET.iterparse(filepath) for _, el in it: _, _, el.tag = el.tag.rpartition('}') root = it.root # 定位符合条件的outboundRelationship outbound = root.find(".//outboundRelationship[@typeCode='SPRT']") if not outbound: return None # 验证relatedInvestigation下的code节点属性 valid_code = outbound.find(".//relatedInvestigation/code[@code='2' and @codeSystem='2.16.840.1.113883.3.989.2.1.1.22']") if not valid_code: return None # 获取目标属性值 priority_node = outbound.find("priorityNumber") return priority_node.attrib.get(attribute) if priority_node else None
方法2:直接用XPath获取属性值
简化代码逻辑,让XPath直接返回属性值:
import lxml.etree as ET def xml_get_attrib_value(filepath, xpath): it = ET.iterparse(filepath) for _, el in it: _, _, el.tag = el.tag.rpartition('}') root = it.root # 执行XPath直接获取属性值列表,取第一个结果 result = root.xpath(xpath) return result[0] if result else None # 调用示例:传入直接获取属性的XPath target_xpath = ".//outboundRelationship[@typeCode='SPRT'][relatedInvestigation/code[@code='2' and @codeSystem='2.16.840.1.113883.3.989.2.1.1.22']]/priorityNumber/@value" print(xml_get_attrib_value("your_xml_file.xml", target_xpath))
内容的提问来源于stack exchange,提问作者Hemant
相关产品推荐
相关产品推荐

