Python解析带命名空间XML:提取不同层级Value标签值的问题
SolarWinds XML解析:提取所有层级标签值的正确方法
问题场景
给定一段SolarWinds告警规则的XML片段:
<?xml version="1.0" ?> <ArrayOfAlertConditionShelve xmlns="http://schemas.datacontract.org/2004/07/SolarWinds.Orion.Core.Models.Alerting" xmlns:i="http://www.w3.org/2001/XMLSchema-instance"> <AlertConditionShelve> <AndThenTimeInterval i:nil="true"/> <ChainType>Trigger</ChainType> <ConditionTypeID>Core.Dynamic</ConditionTypeID> <Configuration> <AlertConditionDynamic xmlns="http://schemas.datacontract.org/2004/07/SolarWinds.Orion.Core.Alerting.Plugins.Conditions.Dynamic" xmlns:i="http://www.w3.org/2001/XMLSchema-instance"> <ExprTree xmlns:a="http://schemas.datacontract.org/2004/07/SolarWinds.Orion.Core.Models.Alerting"> <a:Child> <a:Expr> <a:Child> <a:Expr> <a:Child/> <a:NodeType>Field</a:NodeType> <a:Value>Orion.NPM.Interfaces|TypeDescription</a:Value> </a:Expr> <a:Expr> <a:Child i:nil="true"/> <a:NodeType>Constant</a:NodeType> <a:Value>IEEE 802.3ad Link Aggregate</a:Value> </a:Expr> </a:Child> <a:NodeType>Operator</a:NodeType> <a:Value>=</a:Value> </a:Expr> </a:Child> </ExprTree> </AlertConditionDynamic> </Configuration> </AlertConditionShelve> </ArrayOfAlertConditionShelve>
原有Python解析代码存在两个问题:输出的=)的
def ParsingXMLviaDOM(xmlstr): xmlparse = xml.dom.minidom.parseString(xmlstr) xmlstr = xmlparse.toprettyxml() conditions = xmlparse.getElementsByTagName("a:Expr") print ' ' for condition in conditions: con = condition.getElementsByTagName("a:Value") print con[0].childNodes[0].nodeValue
问题原因
原有代码通过遍历a:Expr节点,仅提取每个节点下第一个a:Value,但:
getElementsByTagName会递归查找当前节点的所有后代节点,导致子节点的a:Value被重复获取;- 上层运算符节点的
a:Value(如=)并非所在a:Expr的第一个a:Value,因此被遗漏。
解决方案
方案1:直接提取所有标签
通过命名空间定位所有a:Value节点,一次性获取所有层级的
import xml.dom.minidom def ParsingXMLviaDOM(xmlstr): xmlparse = xml.dom.minidom.parseString(xmlstr) # 对应XML中xmlns:a的命名空间 namespace = "http://schemas.datacontract.org/2004/07/SolarWinds.Orion.Core.Models.Alerting" # 获取所有a:Value节点 all_value_nodes = xmlparse.getElementsByTagNameNS(namespace, "Value") for node in all_value_nodes: # 确保节点有文本内容 if node.childNodes: print(node.childNodes[0].nodeValue)
输出结果:
Orion.NPM.Interfaces|TypeDescription IEEE 802.3ad Link Aggregate =
方案2:递归解析表达式结构(更贴合业务逻辑)
如果需要明确条件的层级关系(如字段、运算符、常量的组合),可以递归遍历a:Expr节点,构建完整的条件表达式:
import xml.dom.minidom def parse_expression(expr_node, namespace): # 获取当前节点的类型和值 node_type = expr_node.getElementsByTagNameNS(namespace, "NodeType")[0].childNodes[0].nodeValue value_node = expr_node.getElementsByTagNameNS(namespace, "Value") value = value_node[0].childNodes[0].nodeValue if value_node else "" if node_type == "Operator": # 运算符节点,递归解析左右子表达式 child_exprs = expr_node.getElementsByTagNameNS(namespace, "Child")[0].getElementsByTagNameNS(namespace, "Expr") left_expr = parse_expression(child_exprs[0], namespace) right_expr = parse_expression(child_exprs[1], namespace) return f"({left_expr} {value} {right_expr})" elif node_type in ["Field", "Constant"]: # 字段或常量节点,直接返回值 return value def ParsingXMLviaDOM(xmlstr): xmlparse = xml.dom.minidom.parseString(xmlstr) namespace = "http://schemas.datacontract.org/2004/07/SolarWinds.Orion.Core.Models.Alerting" # 定位根表达式节点 root_expr = xmlparse.getElementsByTagNameNS(namespace, "Expr")[0] # 解析并输出完整条件 print(parse_expression(root_expr, namespace))
输出结果:
(Orion.NPM.Interfaces|TypeDescription = IEEE 802.3ad Link Aggregate)
内容的提问来源于stack exchange,提问作者Fisher Tom
相关产品推荐
相关产品推荐

