如何优化目录遍历代码?现有嵌套实现需精简
目录遍历代码优化方案
问题背景
现有如下目录结构,需要遍历目录,当目录与文件满足特定条件时,对每个加载的文件调用指定函数:
self.directories[0] │ └───maps_dir_1 │ │ file1.json │ │ file2.txt │ │ file3.json │ └───subfolder1 │ │ file111.txt │ │ file112.txt │ │ .. │ └───maps_dir_1 │ │ file4.json │ │ file5.txt │ │ file6.json
当前实现代码嵌套层级过深、可读性差,原代码如下:
for maps_dir in self.directories: for map_dir in os.listdir(maps_dir): if fails_check1(map_dir): continue for filename in os.listdir(os.path.join(maps_dir, map_dir)): if not filename.endswith(".json"): continue file_path = os.path.join(map_dir, filename) if os.path.isfile(file_path): with open(file_path, encoding="utf-8") as demo_json: demo_data: Game = json.load(demo_json) match_id = os.path.splitext(filename)[0] if fails_check2(match_id): continue self.do_stuff(demo_data, match_id)
优化方案
通过路径工具替换、逻辑拆分、扁平化循环三个方向优化,提升代码可读性和可维护性:
1. 依赖导入
使用Python3.4+内置的pathlib模块替代os,路径操作更直观:
from pathlib import Path import json from itertools import chain # 可选,用于链式处理
2. 拆分单一职责函数
将筛选目录、筛选文件、加载处理文件的逻辑拆分为独立函数:
def valid_map_dirs(root_dirs): """生成通过目录校验的map目录路径""" for root_dir in root_dirs: root_path = Path(root_dir) for item in root_path.iterdir(): # 只处理目录,且通过check1校验 if item.is_dir() and not fails_check1(item.name): yield item def valid_json_files(map_dirs): """生成目录下的所有json文件路径""" for map_dir in map_dirs: # 直接匹配当前目录下的json文件,自动忽略子目录 yield from map_dir.glob("*.json") def process_json_file(json_file): """加载json文件,校验match_id,返回数据和ID""" match_id = json_file.stem # 直接获取无扩展名的文件名 if fails_check2(match_id): return None, None # 用pathlib的open方法直接打开文件 with json_file.open(encoding="utf-8") as f: demo_data = json.load(f) return demo_data, match_id
3. 扁平化主流程
主逻辑仅负责串联各函数,层级大幅减少:
# 方式一:分步遍历 for map_dir in valid_map_dirs(self.directories): for json_file in valid_json_files([map_dir]): demo_data, match_id = process_json_file(json_file) if demo_data and match_id: self.do_stuff(demo_data, match_id) # 方式二:链式处理(更紧凑) all_valid_json_files = chain.from_iterable( valid_json_files([map_dir]) for map_dir in valid_map_dirs(self.directories) ) for json_file in all_valid_json_files: demo_data, match_id = process_json_file(json_file) if demo_data and match_id: self.do_stuff(demo_data, match_id)
优化亮点
- 路径操作更简洁:
pathlib的路径对象支持.name、.stem等属性,.glob()直接匹配文件,避免繁琐的字符串拼接和判断。 - 逻辑边界清晰:每个函数只做一件事,后续修改校验规则、文件处理逻辑时,只需调整对应函数,不影响整体流程。
- 嵌套层级降低:原代码3层循环嵌套优化为最多2层,主流程逻辑一目了然。
- 扩展性更强:可在
process_json_file中轻松添加异常捕获(如json解析失败)、日志记录等逻辑,不污染主流程。
内容的提问来源于stack exchange,提问作者J.N.
相关产品推荐
相关产品推荐

