如何用Python遍历多级目录下的文本文件并查找匹配内容?
用Python批量遍历多级目录并查找文件内容
当然可以用Python实现,下面提供两种基于标准库的实用方案,无需额外安装依赖:
方案一:使用 os.walk 遍历目录
os.walk 是Python标准库中专门用于递归遍历目录树的工具,能自动遍历指定根目录下的所有子目录和文件。
import os import re # 替换为你的目标根目录路径 root_directory = "/your/target/directory" # 定义匹配规则:正则表达式支持复杂匹配,简单匹配可直接用字符串 match_pattern = re.compile(r"你要查找的内容") # 递归遍历目录 for current_dir, subdirs, files in os.walk(root_directory): for file_name in files: # 过滤文本文件,可根据需求添加更多后缀(比如.md, .log等) if file_name.endswith((".txt", ".md")): full_file_path = os.path.join(current_dir, file_name) try: # 安全打开文件,指定编码避免乱码 with open(full_file_path, "r", encoding="utf-8") as file: # 若文件过大,建议用下方逐行读取的方式 file_content = file.read() matches = match_pattern.findall(file_content) if matches: print(f"文件 {full_file_path} 中找到匹配内容:{matches}") except Exception as error: print(f"读取文件 {full_file_path} 失败:{str(error)}")
方案二:使用 pathlib 遍历目录(Python 3.4+)
pathlib 是更现代化的路径操作库,语法简洁直观,可读性更强。
from pathlib import Path import re root_directory = Path("/your/target/directory") match_pattern = re.compile(r"你要查找的内容") # 递归查找所有文本文件,*.txt可替换为*后再过滤后缀 for file_path in root_directory.rglob("*.txt"): try: # 直接读取文本内容,无需手动管理文件句柄 file_content = file_path.read_text(encoding="utf-8") matches = match_pattern.findall(file_content) if matches: print(f"文件 {file_path} 中找到匹配内容:{matches}") except Exception as error: print(f"读取文件 {file_path} 失败:{str(error)}")
补充:大文件处理优化
如果目录中有超大文本文件,一次性读取全部内容会占用过多内存,建议逐行读取:
# 以os.walk方案为例,替换读取部分的代码 with open(full_file_path, "r", encoding="utf-8") as file: for line_num, line in enumerate(file, start=1): if match_pattern.search(line): print(f"文件 {full_file_path} 第 {line_num} 行匹配:{line.strip()}")
其他注意事项
- 若文件编码不是UTF-8,可调整
encoding参数(比如encoding="gbk") - 若只需简单字符串匹配,无需正则表达式,直接用
if "目标内容" in file_content即可 - 可将匹配结果写入输出文件,方便后续查看,比如用
with open("result.txt", "a", encoding="utf-8") as f: f.write(...)
内容的提问来源于stack exchange,提问作者Andrés Segarra
相关产品推荐
相关产品推荐

