技术咨询:defusedxml中lxml.objectify.cleanup_namespaces的等价实现及替换lxml.etree.parse以解决Bandit安全扫描报错
Great question! Let's tackle this XML security issue and rewrite your code to use defusedxml properly—since it's the go-to tool for parsing untrusted XML data to block attacks like XXE (XML External Entity) injection.
Why the Original Code Is Risky
lxml.etree.parse doesn't disable dangerous XML features by default when handling untrusted input, which leaves your code exposed to XML injection vulnerabilities. defusedxml patches these gaps while keeping most of lxml's familiar API, making it a safe drop-in replacement.
Replacing objectify.cleanup_namespaces with defusedxml
defusedxml doesn't include a direct equivalent to lxml.objectify or its cleanup_namespaces method, but we can replicate the exact behavior with two straightforward steps:
- Strip namespace prefixes from tags: Your original code already handles this with the loop that removes the
}-delimited prefix from element tags—we'll keep that logic. - Clean up unused namespaces: Use
etree.cleanup_namespaces()from defusedxml's lxml-compatible module. This does the same work asobjectify.cleanup_namespacesby removing orphaned namespace declarations.
Rewritten Code with defusedxml
Here's your updated function, modified to be secure while retaining all original functionality:
from defusedxml.lxml import etree def fn_read_xml_root(xml_file): """ function open xml and remove annotation and return the root node xml_file : xml file to be parsed """ with open(xml_file, "r", encoding="utf-8") as x_file: xml_data = x_file.read() # Use defusedxml's parser for safe parsing parser = etree.XMLParser(remove_blank_text=True) xtree = etree.parse(xml_file, parser) xroot = xtree.getroot() # Strip namespace prefixes from element tags (matches original logic) for elem in xroot.iter(): # Replaced deprecated .getiterator() with modern .iter() if not hasattr(elem.tag, "find"): continue idx = elem.tag.find("}") if idx >= 0: elem.tag = elem.tag[idx + 1:] # Clean up unused namespaces (replaces objectify.deannotate + cleanup_namespaces=True) etree.cleanup_namespaces(xroot) # Return xml data and root node of the file return xml_data, xroot
Key Changes Breakdown:
- Safe Parsing: We import
etreefromdefusedxml.lxmlinstead of rawlxml, soetree.parseautomatically blocks XML attacks. - Modernized Iteration: Swapped
.getiterator()for.iter()—the recommended, non-deprecated method in lxml/defusedxml. - Namespace Cleanup: Replaced
objectify.deannotate(cleanup_namespaces=True)withetree.cleanup_namespaces(xroot)to remove unused namespace declarations, matching the original behavior exactly.
This version fixes the Bandit security warning while keeping all your original code's functionality intact.
内容的提问来源于stack exchange,提问作者Ankit Gupta

