Python提取Excel(XLSX)中connections.xml的<dbPr>节点遇问题求助
问题修复方案
核心问题分析
- 节点遍历逻辑错误:XML解析后,
childNodes会包含文本节点(比如换行、空格),原代码直接遍历所有子节点,导致后续处理了无效节点,且没有直接定位到目标<dbPr>节点。 - 属性访问方式错误:
attributes.values()返回的是dict_values对象,不支持下标[0]/[1]访问,这是触发TypeError: 'dict_values' object is not subscriptable的直接原因。
修正后的代码
import zipfile from xml.dom.minidom import parseString def writeoutput(filename, dsn, sql): # 可根据需求修改输出逻辑,比如写入文件 print(f"文件路径: {filename}") print(f"连接字符串: {dsn}") print(f"SQL语句: {sql}\n") def checkfile(filename): if zipfile.is_zipfile(filename): with zipfile.ZipFile(filename, 'r') as zf: if "xl/connections.xml" in zf.namelist(): print(f"检测到连接配置文件: {filename}") xml_content = zf.read('xl/connections.xml') root = parseString(xml_content) # 获取所有connection节点 connections = root.getElementsByTagName('connection') for con in connections: # 直接定位目标dbPr节点,跳过无效文本节点 db_pr_nodes = con.getElementsByTagName('dbPr') if db_pr_nodes: db_pr = db_pr_nodes[0] # 通过属性名直接取值,避免下标访问错误 dsn = db_pr.getAttribute('connection') sql = db_pr.getAttribute('command') writeoutput(filename, dsn, sql)
关键修复点说明
- 精准定位节点:用
getElementsByTagName('dbPr')直接从<connection>节点下获取目标节点,无需遍历所有子节点,避免处理无效的文本节点。 - 属性访问优化:替换
attributes.values()下标访问为getAttribute('属性名')方法,既解决了类型错误,又让代码可读性、健壮性更强。 - 资源安全管理:使用
with语句管理ZipFile对象,确保文件自动关闭,避免资源泄漏。 - 移除冗余异常捕获:原代码的
try-except会掩盖真实错误,修复后可精准排查问题,如需捕获异常可针对性添加。
内容的提问来源于stack exchange,提问作者Fandango68
相关产品推荐
相关产品推荐

