You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python遍历多级目录下的文本文件并查找匹配内容?

用Python批量遍历多级目录并查找文件内容

当然可以用Python实现,下面提供两种基于标准库的实用方案,无需额外安装依赖:

方案一:使用 os.walk 遍历目录

os.walk 是Python标准库中专门用于递归遍历目录树的工具,能自动遍历指定根目录下的所有子目录和文件。

import os
import re

# 替换为你的目标根目录路径
root_directory = "/your/target/directory"
# 定义匹配规则:正则表达式支持复杂匹配,简单匹配可直接用字符串
match_pattern = re.compile(r"你要查找的内容")

# 递归遍历目录
for current_dir, subdirs, files in os.walk(root_directory):
    for file_name in files:
        # 过滤文本文件,可根据需求添加更多后缀(比如.md, .log等)
        if file_name.endswith((".txt", ".md")):
            full_file_path = os.path.join(current_dir, file_name)
            try:
                # 安全打开文件,指定编码避免乱码
                with open(full_file_path, "r", encoding="utf-8") as file:
                    # 若文件过大,建议用下方逐行读取的方式
                    file_content = file.read()
                    matches = match_pattern.findall(file_content)
                    if matches:
                        print(f"文件 {full_file_path} 中找到匹配内容:{matches}")
            except Exception as error:
                print(f"读取文件 {full_file_path} 失败:{str(error)}")

方案二:使用 pathlib 遍历目录(Python 3.4+)

pathlib 是更现代化的路径操作库,语法简洁直观,可读性更强。

from pathlib import Path
import re

root_directory = Path("/your/target/directory")
match_pattern = re.compile(r"你要查找的内容")

# 递归查找所有文本文件,*.txt可替换为*后再过滤后缀
for file_path in root_directory.rglob("*.txt"):
    try:
        # 直接读取文本内容,无需手动管理文件句柄
        file_content = file_path.read_text(encoding="utf-8")
        matches = match_pattern.findall(file_content)
        if matches:
            print(f"文件 {file_path} 中找到匹配内容:{matches}")
    except Exception as error:
        print(f"读取文件 {file_path} 失败:{str(error)}")

补充:大文件处理优化

如果目录中有超大文本文件,一次性读取全部内容会占用过多内存,建议逐行读取:

# 以os.walk方案为例,替换读取部分的代码
with open(full_file_path, "r", encoding="utf-8") as file:
    for line_num, line in enumerate(file, start=1):
        if match_pattern.search(line):
            print(f"文件 {full_file_path} 第 {line_num} 行匹配:{line.strip()}")

其他注意事项

  • 若文件编码不是UTF-8,可调整encoding参数(比如encoding="gbk")
  • 若只需简单字符串匹配,无需正则表达式,直接用if "目标内容" in file_content即可
  • 可将匹配结果写入输出文件,方便后续查看,比如用with open("result.txt", "a", encoding="utf-8") as f: f.write(...)

内容的提问来源于stack exchange,提问作者Andrés Segarra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 15:20:25