You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用listdir()处理文件遇子目录报错,求os.walk()实现方案

问题分析

原代码用os.listdir()遍历目录时,会把子目录本身当成文件去打开,这直接导致了UnicodeDecodeError(目录不是文本文件,没法用UTF-8解码)。另外原代码既没递归处理子目录里的XML文件,也没做编码异常防护,一旦遇到非UTF-8编码的文件就会崩溃。

修复后的代码
from bs4 import BeautifulSoup
import os

input_path = "./input-dir"
output_path = "./output-dir"

# 读取需添加前缀的类名列表,简化处理逻辑
with open("classes.txt", "r", encoding="utf-8") as cls:
    clss = [line.strip() for line in cls.readlines()]

# 用os.walk递归遍历所有层级的目录
for root, dirs, files in os.walk(input_path):
    # 只处理当前目录下的文件
    for filename in files:
        # 过滤出XML文件(可根据实际需求调整后缀)
        if not filename.endswith(".xml"):
            continue
        
        # 构建输入文件的完整路径
        input_file = os.path.join(root, filename)
        # 保持输出目录和输入目录结构一致
        relative_path = os.path.relpath(root, input_path)
        output_dir = os.path.join(output_path, relative_path)
        os.makedirs(output_dir, exist_ok=True)  # 不存在则创建目录,已存在则跳过
        output_file = os.path.join(output_dir, filename)
        
        try:
            # 读取文件,指定编码并处理编码异常
            with open(input_file, "r", encoding="utf-8", errors="replace") as file:
                content = file.read()
            
            # 处理XML内容
            bs_content = BeautifulSoup(content, "lxml")
            str_bs_content = str(bs_content)
            str_bs_content = str_bs_content.replace('<?xml version="1.0" encoding="UTF-8"?><html><body>', "")
            str_bs_content = str_bs_content.replace("</body></html>", "")
            
            # 为目标类名添加前缀
            for cls_name in clss:
                str_bs_content = str_bs_content.replace(cls_name, f"prefix-{cls_name}")
            
            # 写入处理后的内容到输出文件
            with open(output_file, "w", encoding="utf-8") as f:
                f.write(str_bs_content)
        
        except Exception as e:
            print(f"处理文件 {input_file} 时出错: {str(e)}")
关键修改点
  • 替换os.listdir()为os.walk():自动递归遍历所有子目录,区分目录和文件,彻底避免把目录当文件打开的错误。
  • 保留目录结构:输出文件和输入文件的层级结构完全一致,不会把所有文件都堆在output-dir根目录。
  • 过滤目标文件:只处理XML后缀的文件,跳过其他无关文件。
  • 编码异常防护:打开文件时指定UTF-8编码,同时用errors="replace"处理编码不兼容的情况,避免程序崩溃。
  • 简化代码逻辑:用列表推导式处理类名列表,去掉冗余的索引遍历,代码更简洁易读。

内容的提问来源于stack exchange,提问作者Madamadam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 10:40:37