使用pathlib递归检查文件名前缀与父目录前缀匹配问题排查
问题排查与修复
1. 生成器被耗尽导致循环未执行
你通过p = Path(...).rglob('*')创建的是生成器对象,这类对象只能被遍历一次。之前已经用filePaths = [x for x in p if x.is_file()]把生成器的内容全部取完了,后续再执行for element in p时,生成器已经空了,循环根本不会运行,自然不会检测到任何不匹配的文件。
2. Path对象属性调用错误
代码里的element.Path.name和element.Path.parent是错误写法:
- Path实例本身就自带
name(文件名)和parent(父目录对象)属性,正确写法是element.name和element.parent.name - 这种错误会引发
AttributeError,只是因为生成器耗尽导致循环没执行,所以你没看到报错
3. 其他潜在问题
- 未过滤目录:
rglob('*')会返回目录和文件,你需要只检查文件,否则可能把目录也纳入检测 - 前缀匹配逻辑的严谨性:如果文件名或目录名不含下划线,
split("_")[0:1]会返回整个名称的列表,可能不符合你的预期;另外你需求是文件名以父目录前缀(如abc2022_)开头,当前代码是分割后对比前部分,和需求有细微差异
修复后的完整代码
from pathlib import Path # 读取文件列表,用with语句自动关闭文件更安全 with open("fileList.txt", "r") as fileList: data = fileList.read() fileList_reformatted = data.replace('\n', '').split(",") print(fileList_reformatted) # 获取目标目录下所有文件,直接转为列表避免生成器耗尽问题 target_dir = Path('C:/Users/Common/Downloads/compare') filePaths = [x for x in target_dir.rglob('*') if x.is_file()] filePaths_string = [str(x) for x in filePaths] print(filePaths_string) # 查找缺失的预期文件 differences1 = [] for element in fileList_reformatted: if element not in filePaths_string: differences1.append(element) print("The following files from the provided list were not found:", differences1) # 查找多余的意外文件 differences2 = [] for element in filePaths_string: if element not in fileList_reformatted: differences2.append(element) print("The following unexpected files were found:", differences2) # 检查文件名前缀与父目录前缀匹配 wrong_location = [] for element in filePaths: parent_name = element.parent.name # 提取父目录前缀(比如abc2022_001 -> abc2022_) if '_' in parent_name: parent_prefix = parent_name.split('_')[0] + '_' else: # 若父目录无下划线,可根据需求调整逻辑,比如用整个目录名加下划线 parent_prefix = parent_name + '_' # 检查文件名是否以父目录前缀开头 if not element.name.startswith(parent_prefix): wrong_location.append(str(element)) print("Following files may be in the wrong location:", wrong_location)
内容的提问来源于stack exchange,提问作者Paul
相关产品推荐
相关产品推荐

