You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历子目录XML文件报错:TypeError路径对象类型不匹配

问题描述

我是Python新手,目前遇到无法解决的问题:想要遍历子目录中的XML文件,打开并读取其内容。

原本遍历子目录的代码可以正常运行:

record_count = 0
for root, dirs, files in os.walk('/testing/'):
    for file in files:
        if file.endswith('.xml'):
            record_count = record_count + 1
            path = os.path.join(root, file)
            print(record_count, path)

输出符合预期,path变量每次迭代都会正确更新:

1 /testing/subfolder/mufs190-1951-dpl2k4-i001.reference_export.xml
2 /testing/subfolder/mufs190-1951-dpl2k3-i001.reference_export.xml

但添加读取XML文件并提取信息的代码后:

record_count = 0
for root, dirs, files in os.walk('/testing/'):
    for file in files:
        if file.endswith('.xml'):
            record_count = record_count + 1
            path = os.path.join(root, file)
            print(record_count, path)

            with open(path, 'r') as f:
                tree = ET.parse(f)
                root = tree.getroot()
                # Retrieve the item's title
                for citeinfo in root.findall('./idinfo/citation/citeinfo'):
                    title = citeinfo.find('title').text
                print(title)

代码首次迭代成功,之后触发错误:

Traceback (most recent call last):
  File "/testing/subfolder.py", line 25, in <module>
    path = os.path.join(root, file)
  File "/usr/local/Cellar/python@3.9/3.9.6/Frameworks/Python.framework/Versions/3.9/lib/python3.9/posixpath.py", line 76, in join
    a = os.fspath(a)
TypeError: expected str, bytes or os.PathLike object, not Element
问题原因与解决方法

问题出在变量名冲突:你在os.walk循环里用root作为存储目录路径的变量,之后解析XML时又执行root = tree.getroot(),把这个变量重新赋值成了XML根节点对象。

第一次迭代后,原本是字符串类型的目录路径变量root被覆盖成Element类型的XML节点,第二次循环执行os.path.join(root, file)时,自然会报错——因为os.path.join需要字符串路径,而非XML节点对象。

修正方法很简单:把XML解析部分的root变量改成其他名字,比如xml_root:

import os
import xml.etree.ElementTree as ET

record_count = 0
for root, dirs, files in os.walk('/testing/'):
    for file in files:
        if file.endswith('.xml'):
            record_count = record_count + 1
            path = os.path.join(root, file)
            print(record_count, path)

            with open(path, 'r') as f:
                tree = ET.parse(f)
                xml_root = tree.getroot()  # 修改变量名避免冲突
                # Retrieve the item's title
                for citeinfo in xml_root.findall('./idinfo/citation/citeinfo'):
                    title = citeinfo.find('title').text if citeinfo.find('title') is not None else '无标题'
                print(title)

额外补充两个实用细节:

  • 代码开头要导入xml.etree.ElementTree模块(你的代码里用了ET别名,所以需要加import xml.etree.ElementTree as ET)
  • 增加title节点的存在判断,避免部分XML文件没有该节点时触发报错

内容的提问来源于stack exchange,提问作者spacelibrarian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 03:33:28