如何在文件夹及其子文件夹的TXT文件中搜索指定字符串并列出文件
遍历文件夹搜索TXT文件中的指定字符串
你原来的open('*', 'r').read().find('Test')写法行不通,因为open()函数不支持直接用通配符*批量打开文件,得先递归遍历目标文件夹下的所有TXT文件,再逐个检查内容是否包含指定字符串。
下面是两种可行的实现方式:
方式一:使用os模块(兼容Python2/3)
import os def search_txt_files(root_dir, target_str): # 存储包含目标字符串的文件路径 matched_files = [] # 递归遍历文件夹 for dirpath, _, filenames in os.walk(root_dir): for filename in filenames: # 筛选TXT文件 if filename.endswith('.txt'): file_path = os.path.join(dirpath, filename) try: # 读取文件内容并检查 with open(file_path, 'r', encoding='utf-8', errors='ignore') as f: if target_str in f.read(): matched_files.append(file_path) except Exception as e: print(f"读取文件失败:{file_path},错误信息:{str(e)}") return matched_files # 调用示例:替换为你的目标文件夹路径和搜索字符串 result = search_txt_files('./your_target_folder', 'Test') if result: print("包含指定字符串的文件:") for file in result: print(file) else: print("未找到包含指定字符串的TXT文件")
方式二:使用pathlib模块(Python3.4+推荐)
pathlib提供了更简洁的路径操作API,代码可读性更强:
from pathlib import Path def search_txt_files(root_dir, target_str): matched_files = [] # 递归获取所有TXT文件路径 for txt_path in Path(root_dir).rglob('*.txt'): try: with open(txt_path, 'r', encoding='utf-8', errors='ignore') as f: if target_str in f.read(): matched_files.append(str(txt_path)) except Exception as e: print(f"读取文件失败:{txt_path},错误信息:{str(e)}") return matched_files # 调用示例 result = search_txt_files('./your_target_folder', 'Test') if result: print("包含指定字符串的文件:") for file in result: print(file) else: print("未找到包含指定字符串的TXT文件")
注意事项
encoding='utf-8'可以根据你的实际文件编码调整,比如gbk等;errors='ignore'用于跳过编码错误的字符,避免程序崩溃。- 如果文件过大,一次性
read()会占用较多内存,可以改为逐行读取检查:with open(file_path, 'r', encoding='utf-8', errors='ignore') as f: for line in f: if target_str in line: matched_files.append(file_path) break # 找到即停止,提升效率
内容的提问来源于stack exchange,提问作者question12
相关产品推荐
相关产品推荐

