简化XML解析空值判断:避免重复XML路径
解决XML解析中缺失标签的空值处理冗余问题
解析XML时,部分标签非必填,直接调用.text会因标签为None触发报错。当前通过重复调用find()做空值判断的方式冗余度极高,可通过以下几种方案优化:
方法1:封装通用工具函数
写一个统一处理空值逻辑的函数,避免重复编写判断代码:
def get_xml_text(node, xpath, namespaces): elem = node.find(xpath, namespaces) return elem.text.strip() if elem else "" # 调用示例 myArr.append([ get_xml_text(properties, "link[@title='myTitle']/a:inline/feed/entry/content/b:properties/c:Id1", namespaces), get_xml_text(properties, "link[@title='myTitle']/a:inline/feed/entry/content/b:properties/c:Id2", namespaces), get_xml_text(properties, "link[@title='myTitle']/a:inline/feed/entry/content/b:properties/c:Id3", namespaces), # ... 其他字段 ])
方法2:用getattr()简化一行判断
利用Python内置的getattr()指定默认值,一行完成空值处理:
myArr.append([ getattr(properties.find("link[@title='myTitle']/a:inline/feed/entry/content/b:properties/c:Id1", namespaces), 'text', "").strip(), getattr(properties.find("link[@title='myTitle']/a:inline/feed/entry/content/b:properties/c:Id2", namespaces), 'text', "").strip(), # ... 其他字段 ])
如果find()返回None,getattr会直接返回空字符串,避免报错。
方法3:提前获取公共父节点减少重复路径
观察到所有字段的路径前缀一致,先获取公共父节点,后续仅查找子节点,既减少冗余又提升效率:
# 先获取公共父节点 parent_node = properties.find("link[@title='myTitle']/a:inline/feed/entry/content/b:properties", namespaces) def get_child_text(parent, child_xpath): if not parent: return "" elem = parent.find(child_xpath) return elem.text.strip() if elem else "" # 调用示例 myArr.append([ get_child_text(parent_node, "c:Id1"), get_child_text(parent_node, "c:Id2"), get_child_text(parent_node, "c:Id3"), # ... 其他字段 ])
内容的提问来源于stack exchange,提问作者JAM
相关产品推荐
相关产品推荐

