You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术咨询:defusedxml中lxml.objectify.cleanup_namespaces的等价实现及替换lxml.etree.parse以解决Bandit安全扫描报错

Great question! Let's tackle this XML security issue and rewrite your code to use defusedxml properly—since it's the go-to tool for parsing untrusted XML data to block attacks like XXE (XML External Entity) injection.

Why the Original Code Is Risky

lxml.etree.parse doesn't disable dangerous XML features by default when handling untrusted input, which leaves your code exposed to XML injection vulnerabilities. defusedxml patches these gaps while keeping most of lxml's familiar API, making it a safe drop-in replacement.

Replacing objectify.cleanup_namespaces with defusedxml

defusedxml doesn't include a direct equivalent to lxml.objectify or its cleanup_namespaces method, but we can replicate the exact behavior with two straightforward steps:

  1. Strip namespace prefixes from tags: Your original code already handles this with the loop that removes the }-delimited prefix from element tags—we'll keep that logic.
  2. Clean up unused namespaces: Use etree.cleanup_namespaces() from defusedxml's lxml-compatible module. This does the same work as objectify.cleanup_namespaces by removing orphaned namespace declarations.

Rewritten Code with defusedxml

Here's your updated function, modified to be secure while retaining all original functionality:

from defusedxml.lxml import etree

def fn_read_xml_root(xml_file):
    """ function open xml and remove annotation and return the root node
    xml_file : xml file to be parsed
    """
    with open(xml_file, "r", encoding="utf-8") as x_file:
        xml_data = x_file.read()
    
    # Use defusedxml's parser for safe parsing
    parser = etree.XMLParser(remove_blank_text=True)
    xtree = etree.parse(xml_file, parser)
    xroot = xtree.getroot()
    
    # Strip namespace prefixes from element tags (matches original logic)
    for elem in xroot.iter():  # Replaced deprecated .getiterator() with modern .iter()
        if not hasattr(elem.tag, "find"):
            continue
        idx = elem.tag.find("}")
        if idx >= 0:
            elem.tag = elem.tag[idx + 1:]
    
    # Clean up unused namespaces (replaces objectify.deannotate + cleanup_namespaces=True)
    etree.cleanup_namespaces(xroot)
    
    # Return xml data and root node of the file
    return xml_data, xroot

Key Changes Breakdown:

  • Safe Parsing: We import etree from defusedxml.lxml instead of raw lxml, so etree.parse automatically blocks XML attacks.
  • Modernized Iteration: Swapped .getiterator() for .iter()—the recommended, non-deprecated method in lxml/defusedxml.
  • Namespace Cleanup: Replaced objectify.deannotate(cleanup_namespaces=True) with etree.cleanup_namespaces(xroot) to remove unused namespace declarations, matching the original behavior exactly.

This version fixes the Bandit security warning while keeping all your original code's functionality intact.

内容的提问来源于stack exchange,提问作者Ankit Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 13:52:35