使用listdir()处理文件遇子目录报错,求os.walk()实现方案
问题分析
原代码用os.listdir()遍历目录时,会把子目录本身当成文件去打开,这直接导致了UnicodeDecodeError(目录不是文本文件,没法用UTF-8解码)。另外原代码既没递归处理子目录里的XML文件,也没做编码异常防护,一旦遇到非UTF-8编码的文件就会崩溃。
修复后的代码
from bs4 import BeautifulSoup import os input_path = "./input-dir" output_path = "./output-dir" # 读取需添加前缀的类名列表,简化处理逻辑 with open("classes.txt", "r", encoding="utf-8") as cls: clss = [line.strip() for line in cls.readlines()] # 用os.walk递归遍历所有层级的目录 for root, dirs, files in os.walk(input_path): # 只处理当前目录下的文件 for filename in files: # 过滤出XML文件(可根据实际需求调整后缀) if not filename.endswith(".xml"): continue # 构建输入文件的完整路径 input_file = os.path.join(root, filename) # 保持输出目录和输入目录结构一致 relative_path = os.path.relpath(root, input_path) output_dir = os.path.join(output_path, relative_path) os.makedirs(output_dir, exist_ok=True) # 不存在则创建目录,已存在则跳过 output_file = os.path.join(output_dir, filename) try: # 读取文件,指定编码并处理编码异常 with open(input_file, "r", encoding="utf-8", errors="replace") as file: content = file.read() # 处理XML内容 bs_content = BeautifulSoup(content, "lxml") str_bs_content = str(bs_content) str_bs_content = str_bs_content.replace('<?xml version="1.0" encoding="UTF-8"?><html><body>', "") str_bs_content = str_bs_content.replace("</body></html>", "") # 为目标类名添加前缀 for cls_name in clss: str_bs_content = str_bs_content.replace(cls_name, f"prefix-{cls_name}") # 写入处理后的内容到输出文件 with open(output_file, "w", encoding="utf-8") as f: f.write(str_bs_content) except Exception as e: print(f"处理文件 {input_file} 时出错: {str(e)}")
关键修改点
- 替换
os.listdir()为os.walk():自动递归遍历所有子目录,区分目录和文件,彻底避免把目录当文件打开的错误。 - 保留目录结构:输出文件和输入文件的层级结构完全一致,不会把所有文件都堆在output-dir根目录。
- 过滤目标文件:只处理XML后缀的文件,跳过其他无关文件。
- 编码异常防护:打开文件时指定UTF-8编码,同时用
errors="replace"处理编码不兼容的情况,避免程序崩溃。 - 简化代码逻辑:用列表推导式处理类名列表,去掉冗余的索引遍历,代码更简洁易读。
内容的提问来源于stack exchange,提问作者Madamadam
相关产品推荐
相关产品推荐

