超大型文件目录下高效获取JSON文件路径的技术方案咨询
高效遍历超大型目录获取所有JSON文件路径的方案
针对超大型目录(数百万文件、数万个文件夹)的场景,glob.glob因性能瓶颈无法满足需求,以下是几种更高效的解决方案:
方案1:使用os.scandir(Python原生高效API)
os.scandir是Python 3.5+引入的文件遍历API,比glob和os.listdir性能提升显著——它返回的DirEntry对象缓存了文件属性(如是否为目录),大幅减少系统调用次数,同时迭代式遍历避免了内存过载和递归栈溢出问题。
import os def find_json_files(root_dir): json_paths = [] # 用栈实现迭代式遍历,规避递归深度限制 stack = [root_dir] while stack: current_dir = stack.pop() with os.scandir(current_dir) as entries: for entry in entries: if entry.is_file(follow_symlinks=False) and entry.name.endswith('.json'): json_paths.append(entry.path) elif entry.is_dir(follow_symlinks=False): stack.append(entry.path) return json_paths # 使用示例 directory = "你的根目录路径" json_files = find_json_files(directory)
方案2:调用系统原生命令(性能最优)
系统原生的文件查找命令由C实现,直接操作文件系统,无Python层额外开销,是超大规模场景下最快的选择:
Windows系统
import subprocess def find_json_windows(root_dir): result = subprocess.run( ['dir', root_dir, '/s', '/b', '*.json'], capture_output=True, text=True, shell=True ) if result.returncode == 0: return result.stdout.strip().split('\n') else: raise RuntimeError(f"命令执行失败: {result.stderr}") # 使用示例 directory = "你的根目录路径" json_files = find_json_windows(directory)
Linux/macOS系统
import subprocess def find_json_unix(root_dir): result = subprocess.run( ['find', root_dir, '-type', 'f', '-name', '*.json'], capture_output=True, text=True ) if result.returncode == 0: return result.stdout.strip().split('\n') else: raise RuntimeError(f"命令执行失败: {result.stderr}") # 使用示例 directory = "/你的根目录路径" json_files = find_json_unix(directory)
方案3:使用pathlib.rglob(简洁高效的替代方案)
pathlib.rglob底层基于os.scandir实现,性能优于原生glob.glob,同时代码更简洁:
from pathlib import Path def find_json_pathlib(root_dir): return [str(path) for path in Path(root_dir).rglob('*.json')] # 使用示例 directory = "你的根目录路径" json_files = find_json_pathlib(directory)
内容的提问来源于stack exchange,提问作者RagingMonkey
相关产品推荐
相关产品推荐

