如何优化Python目录生成脚本:仅遍历一次文件完成分类
优化VS Code Notes扩展目录生成脚本的效率问题
原脚本的get_body()函数会针对每个分类重复扫描所有文件,效率极低。以下是优化后的版本,仅需遍历一次文件即可完成分类与目录生成:
优化后的完整脚本
""" Python script to generate a table of contents .md file for VSCode "Notes" users Run script in Notes.notesLocation to generate, then open _toc.md in preview mode User must prefix note file names with corresponding values in cats (categories) i.e. dj_admin_model.md, py_polymorphism.md, st_ascii.md User could put a top link in every note to quickly return to table of contents e.g. [< content](_toc.md) """ import os dbug = True path = '.' ftyp = '.md' file = '_toc.md' cats = { 'Config' : '_', 'Django' : 'dj_', 'Markdown' : 'md_', 'Python' : 'py_', 'Standard' : 'st_', 'VSCode' : 'vs_', } def get_files(): for _, _, files in os.walk(path): return (f for f in files if f.lower().endswith(ftyp.lower())) def get_body(): # 初始化分类文件字典,每个分类对应空列表 category_files = {key: [] for key in cats.keys()} # 仅遍历一次所有文件,分配到对应分类 for f in get_files(): matched = False for cat_name, prefix in cats.items(): if f.startswith(prefix): # 提取文件显示名称并添加到分类列表 display_name = f.replace('.md', '').split('_', 1)[1] if '_' in f else f.replace('.md', '') category_files[cat_name].append(f"- [{display_name}]({f})\n") matched = True break # 匹配到分类后跳出,避免重复判断 # 生成最终目录内容 body = f'["{file}" generated by running "{__file__}".]: #\n\n# Content\n\n' for cat_name, items in category_files.items(): body += f'### {cat_name}\n' body += ''.join(items) body += '\n' return body def write_toc(): with open(file, mode='wt') as f: f.write(get_body()) def print_toc(): with open(file) as f: print(f.read()) def main(): write_toc() if dbug: print('_'*60) print_toc() if __name__ == '__main__': main()
优化说明
- 单次文件扫描:仅调用一次
get_files()获取所有目标文件,避免重复遍历文件系统 - 预分组文件:通过
category_files字典提前将文件分配到对应分类,仅遍历文件一次 - 输出格式兼容:生成的
_toc.md内容与原脚本完全一致,不影响原有使用逻辑
内容的提问来源于stack exchange,提问作者Mlek D. Nairb
相关产品推荐
相关产品推荐

